A generalist world model conditioned on action flow: robot actions represented as pixel motion. At deployment it runs as a hybrid simulator: </p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/69e9919b48b74358bf7d4dab/qeBMPFBj039nupCJFOmHz.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\na physics engine moves the robot, a learned video model predicts what the world does in response. One model trains across embodiments, simulates them, evaluates policies, and drives a real robot.","updatedAt":"2026-08-24T14:50:31.850Z","author":{"_id":"69e9919b48b74358bf7d4dab","avatarUrl":"/avatars/cb1ae4f6faab74732b68dd647d6754d3.svg","fullname":"Hongyu Li","name":"hongyu-nvidia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8717257976531982},"editors":["hongyu-nvidia"],"editorAvatarUrls":["/avatars/cb1ae4f6faab74732b68dd647d6754d3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.18077","authors":[{"_id":"6a878b6c89e517cbfd75db70","user":{"_id":"69e9919b48b74358bf7d4dab","avatarUrl":"/avatars/cb1ae4f6faab74732b68dd647d6754d3.svg","isPro":false,"fullname":"Hongyu Li","user":"hongyu-nvidia","type":"user","name":"hongyu-nvidia"},"name":"Hongyu Li","status":"claimed_verified","statusLastChangedAt":"2026-08-21T08:45:05.175Z","hidden":false},{"_id":"6a878b6c89e517cbfd75db71","name":"Bowen Wen","hidden":false},{"_id":"6a878b6c89e517cbfd75db72","name":"Xinghao Zhu","hidden":false},{"_id":"6a878b6c89e517cbfd75db73","name":"Yixuan Wang","hidden":false},{"_id":"6a878b6c89e517cbfd75db74","name":"Yilun Du","hidden":false},{"_id":"6a878b6c89e517cbfd75db75","name":"Yunzhu Li","hidden":false},{"_id":"6a878b6c89e517cbfd75db76","name":"George Konidaris","hidden":false},{"_id":"6a878b6c89e517cbfd75db77","name":"Stan Birchfield","hidden":false},{"_id":"6a878b6c89e517cbfd75db78","name":"Soha Pouya","hidden":false},{"_id":"6a878b6c89e517cbfd75db79","name":"Chenran Li","hidden":false},{"_id":"6a878b6c89e517cbfd75db7a","name":"Yan Chang","hidden":false}],"publishedAt":"2026-08-18T17:59:30.000Z","submittedOnDailyAt":"2026-08-24T00:00:00.000Z","title":"Hydra-0: Action Flow for Generalist World Modeling and Control","submittedOnDailyBy":{"_id":"69e9919b48b74358bf7d4dab","avatarUrl":"/avatars/cb1ae4f6faab74732b68dd647d6754d3.svg","isPro":false,"fullname":"Hongyu Li","user":"hongyu-nvidia","type":"user","name":"hongyu-nvidia"},"summary":"We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.","upvotes":2,"discussionId":"6a878b6c89e517cbfd75db7b","projectPage":"https://nvidia-isaac.github.io/video_to_data/hydra-0/","ai_summary":"Hydra-0 uses action flow as a shared visual interface for generalist world modeling and robot control across diverse embodiments and tasks.","ai_keywords":["world model","action flow","pixel motion","cross-embodiment","zero-shot composition","inverse mode","world action model","open-loop policy evaluation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69e9919b48b74358bf7d4dab","avatarUrl":"/avatars/cb1ae4f6faab74732b68dd647d6754d3.svg","isPro":false,"fullname":"Hongyu Li","user":"hongyu-nvidia","type":"user"},{"_id":"69a5cba5ee290d6bb49457b8","avatarUrl":"/avatars/f80c17c13d6baf6bcd375d31efe21116.svg","isPro":false,"fullname":"Darrow O'Lykos","user":"darrowoflykos","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.18077.md","query":{}}">
Hydra-0: Action Flow for Generalist World Modeling and Control
Abstract
Hydra-0 uses action flow as a shared visual interface for generalist world modeling and robot control across diverse embodiments and tasks.
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.
Community
A generalist world model conditioned on action flow: robot actions represented as pixel motion. At deployment it runs as a hybrid simulator:
a physics engine moves the robot, a learned video model predicts what the world does in response. One model trains across embodiments, simulates them, evaluates policies, and drives a real robot.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.18077 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.18077 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.18077 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.