Hugging Face Daily Papers · · 4 min read

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

At the time of writing, DreamX-Phi 1.0 achieves first place on Track 1 and second place on Track 2 of the WorldArena 2.0 Challenge. Model weights and inference code will be made publicly available (<a href=\"https://github.com/AMAP-ML/DreamX-Phi\" rel=\"nofollow\">https://github.com/AMAP-ML/DreamX-Phi</a>) after the WorldArena 2.0 IROS Challenge concludes.</p>\n","updatedAt":"2026-08-14T03:41:12.005Z","author":{"_id":"66d255e3947594430c723ff6","avatarUrl":"/avatars/c56e4792332a01bf34085a75ee64916e.svg","fullname":"xiaochonglinghu","name":"xiaochonglinghu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":2,"identifiedLanguage":{"language":"en","probability":0.8229163885116577},"editors":["xiaochonglinghu"],"editorAvatarUrls":["/avatars/c56e4792332a01bf34085a75ee64916e.svg"],"reactions":[{"reaction":"🚀","users":["ruichen9618"],"count":1},{"reaction":"🔥","users":["ruichen9618"],"count":1},{"reaction":"👀","users":["ruichen9618"],"count":1}],"isReport":false}},{"id":"6a7e8cb7c103b83b6cd15ed9","author":{"_id":"66d255e3947594430c723ff6","avatarUrl":"/avatars/c56e4792332a01bf34085a75ee64916e.svg","fullname":"xiaochonglinghu","name":"xiaochonglinghu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false},"createdAt":"2026-08-14T03:34:15.000Z","type":"comment","data":{"edited":true,"hidden":true,"hiddenBy":"","hiddenReason":"Resolved","latest":{"raw":"This comment has been hidden","html":"This comment has been hidden","updatedAt":"2026-08-14T03:35:02.919Z","author":{"_id":"66d255e3947594430c723ff6","avatarUrl":"/avatars/c56e4792332a01bf34085a75ee64916e.svg","fullname":"xiaochonglinghu","name":"xiaochonglinghu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"editors":[],"editorAvatarUrls":[],"reactions":[]}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.13489","authors":[{"_id":"6a7e820342823931a1f17673","name":"DreamX Team","hidden":false},{"_id":"6a7e820342823931a1f17674","user":{"_id":"64be128a2e66dc7b8bd8459d","avatarUrl":"/avatars/ac5a5246dc19dd35bbd89d7fc492cba5.svg","isPro":false,"fullname":"Rui Chen","user":"ruichen9618","type":"user","name":"ruichen9618"},"name":"Rui Chen","status":"claimed_verified","statusLastChangedAt":"2026-08-14T08:45:04.692Z","hidden":false},{"_id":"6a7e820342823931a1f17675","name":"Xiangxiang Chu","hidden":false},{"_id":"6a7e820342823931a1f17676","name":"Geng Li","hidden":false},{"_id":"6a7e820342823931a1f17677","name":"Jifan Li","hidden":false},{"_id":"6a7e820342823931a1f17678","name":"Qingfeng Shi","hidden":false},{"_id":"6a7e820342823931a1f17679","name":"Datao Tang","hidden":false},{"_id":"6a7e820342823931a1f1767a","name":"Jing Tang","hidden":false},{"_id":"6a7e820342823931a1f1767b","name":"Jun Wang","hidden":false},{"_id":"6a7e820342823931a1f1767c","name":"Pengfei Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66d255e3947594430c723ff6/A5LjPr64bqOeWiNUHMxEJ.mp4"],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation","submittedOnDailyBy":{"_id":"66d255e3947594430c723ff6","avatarUrl":"/avatars/c56e4792332a01bf34085a75ee64916e.svg","isPro":false,"fullname":"xiaochonglinghu","user":"xiaochonglinghu","type":"user","name":"xiaochonglinghu"},"summary":"We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via PRoPE-style geometric encoding, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight depth branch for scene-level geometry and use SAM3 masks with a frozen V-JEPA teacher to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.","upvotes":76,"discussionId":"6a7e820442823931a1f1767d","githubRepo":"https://github.com/AMAP-ML/DreamX-Phi","githubRepoAddedBy":"user","ai_summary":"DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that uses geometric attention encoding, depth estimation, object masks with a frozen teacher, and distillation to generate faithful future observations.","ai_keywords":["action-conditioned video world model","SE(3) transformations","PRoPE-style geometric encoding","depth branch","SAM3 masks","V-JEPA teacher","distribution-matching distillation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":23,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66d255e3947594430c723ff6","avatarUrl":"/avatars/c56e4792332a01bf34085a75ee64916e.svg","isPro":false,"fullname":"xiaochonglinghu","user":"xiaochonglinghu","type":"user"},{"_id":"64d1dc5273174cecdffc97d3","avatarUrl":"/avatars/6564e6b68fee9673f75b6366adf39a3b.svg","isPro":false,"fullname":"Wang Yong","user":"seashell11","type":"user"},{"_id":"668b695b724d9307437c6995","avatarUrl":"/avatars/f0f9ac4ed79de40216bbcb983b10e2a2.svg","isPro":false,"fullname":"Yang Li","user":"yangli2000","type":"user"},{"_id":"661de9defdbc9c247f159d15","avatarUrl":"/avatars/38e21e78327cc908201122405c48f41b.svg","isPro":false,"fullname":"Rui Dai","user":"DerryD","type":"user"},{"_id":"6773bcaa675a971ddf1e81dd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/a8VUwZYXd7O_mq_zFvXMh.png","isPro":false,"fullname":"CokeWang","user":"CokeWang","type":"user"},{"_id":"65003db8bef9b594656f8fa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65003db8bef9b594656f8fa7/L6cvPOAeBRnFnIQwWxYyf.png","isPro":false,"fullname":"Hailang Huang","user":"lerogo","type":"user"},{"_id":"676becb8317c6dbe689474f0","avatarUrl":"/avatars/39209fce8ed644426d36e1a67a192fc2.svg","isPro":false,"fullname":"bill","user":"hunchteller","type":"user"},{"_id":"663bbb61ec1aafe3d6c05558","avatarUrl":"/avatars/b2f593d0ae0adbaad9a7a99b490e3a2b.svg","isPro":false,"fullname":"Ziyu Ma","user":"poiuytrewq123","type":"user"},{"_id":"695b26295e8bad3be68b15f4","avatarUrl":"/avatars/2d4f0988dfa72f3388824ef2f377a546.svg","isPro":false,"fullname":"Fortune Wang","user":"FortuneV2","type":"user"},{"_id":"673c09d251d8d86ed0e4b343","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9p0jpsSw_HYINDHsKasDW.png","isPro":false,"fullname":"guo","user":"sigma28","type":"user"},{"_id":"689e980546d08466f8bc86f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/689e980546d08466f8bc86f8/NdC0BSVmkCjLuyKXRDS9H.jpeg","isPro":false,"fullname":"Xuecai Hu","user":"huxc1208","type":"user"},{"_id":"64906d6d1afdee3acd06ad1a","avatarUrl":"/avatars/4f1978299a93411866f74b4ddd2ef569.svg","isPro":false,"fullname":"Tian Meng","user":"rusuanjun","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.13489.md","query":{}}">
Papers
arxiv:2608.13489

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Published on Aug 13
· Submitted by
xiaochonglinghu
on Aug 14
Authors:
,

Abstract

DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that uses geometric attention encoding, depth estimation, object masks with a frozen teacher, and distillation to generate faithful future observations.

We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via PRoPE-style geometric encoding, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight depth branch for scene-level geometry and use SAM3 masks with a frozen V-JEPA teacher to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.

Community

At the time of writing, DreamX-Phi 1.0 achieves first place on Track 1 and second place on Track 2 of the WorldArena 2.0 Challenge. Model weights and inference code will be made publicly available (https://github.com/AMAP-ML/DreamX-Phi) after the WorldArena 2.0 IROS Challenge concludes.

This comment has been hidden (marked as Resolved)
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.13489
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.13489 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.13489 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.13489 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers