Webiste: <a href=\"https://robbyant-research.github.io/Zero-WAM\" rel=\"nofollow\">https://robbyant-research.github.io/Zero-WAM</a>
<br>Code: <a href=\"https://github.com/robbyant-research/Zero-WAM\" rel=\"nofollow\">https://github.com/robbyant-research/Zero-WAM</a>
<br>Paper: <a href=\"https://arxiv.org/abs/2608.26103\" rel=\"nofollow\">https://arxiv.org/abs/2608.26103</a></p>\n","updatedAt":"2026-08-28T02:38:35.800Z","author":{"_id":"6721acf7bbf9703bc285e840","avatarUrl":"/avatars/487f5cae476758d318465aa9ed103568.svg","fullname":"Yujie Zhao","name":"HomieZ","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5486677289009094},"editors":["HomieZ"],"editorAvatarUrls":["/avatars/487f5cae476758d318465aa9ed103568.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26103","authors":[{"_id":"6a8fb1b02c24e8c5fab329bb","user":{"_id":"6569ca717fd1d421381b1e5a","avatarUrl":"/avatars/080655b2497c2fa44a33da3ccfabb84b.svg","isPro":false,"fullname":"Jiaming Zhou","user":"Jiaming2472","type":"user","name":"Jiaming2472"},"name":"Jiaming Zhou","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.967Z","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329bc","name":"Qihang Zhang","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329bd","name":"Gangwei Xu","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329be","name":"Cunxin Fan","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329bf","name":"Yujie Zhao","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c0","name":"Ruilin Wang","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c1","name":"Yiming Luo","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c2","name":"Shuai Yang","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c3","name":"Xing Zhu","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c4","name":"Yujun Shen","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c5","name":"Junwei Liang","hidden":false},{"_id":"6a8fb1b02c24e8c5fab329c6","name":"Yinghao Xu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6721acf7bbf9703bc285e840/nsQFigd-2V3fayPoW9hsR.mp4"],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-28T00:00:00.000Z","title":"Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization","submittedOnDailyBy":{"_id":"6721acf7bbf9703bc285e840","avatarUrl":"/avatars/487f5cae476758d318465aa9ed103568.svg","isPro":false,"fullname":"Yujie Zhao","user":"HomieZ","type":"user","name":"HomieZ"},"summary":"Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.","upvotes":16,"discussionId":"6a8fb1b02c24e8c5fab329c7","projectPage":"https://robbyant-research.github.io/Zero-WAM","githubRepo":"https://github.com/robbyant-research/Zero-WAM","githubRepoAddedBy":"user","ai_summary":"Zero-WAM enables robotic manipulation of unseen tasks by conditioning a causal video-action model on in-context human video guidance, supported by an automatically generated dataset and a future-chunk prediction objective.","ai_keywords":["zero-shot cross-task generalization","in-context learning","causal video-action model","HumanGen","in-context future chunk prediction","video-action baseline"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":100,"organization":{"_id":"6a8eb8c2b938953de99e7939","name":"Robbyant-Research","fullname":"Robbyant Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/sGSUfQbY1ezHjHO7mgSXL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6721acf7bbf9703bc285e840","avatarUrl":"/avatars/487f5cae476758d318465aa9ed103568.svg","isPro":false,"fullname":"Yujie Zhao","user":"HomieZ","type":"user"},{"_id":"67d4dac5fb7b68b1126ea895","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/eKeHC8RR_GlieyLgg1UZH.png","isPro":false,"fullname":"Minghao Guo","user":"MinghaoGuo","type":"user"},{"_id":"674fca91d9458feb8f3cf7f5","avatarUrl":"/avatars/6d36a6a9d015ae4623b1879a16c9c220.svg","isPro":false,"fullname":"Kun-Yu Lin","user":"LasNack","type":"user"},{"_id":"64291693eb320ead3d2f58a5","avatarUrl":"/avatars/67d593dedcf1f53eac88973d8a8b0fa4.svg","isPro":false,"fullname":"Teli MA","user":"TeliMa","type":"user"},{"_id":"67c0acd7ec3563f6b93d3f51","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/S3eKPZswOffvGVY6U5HIN.png","isPro":false,"fullname":"zifan wang","user":"aCodeDog","type":"user"},{"_id":"6569ca717fd1d421381b1e5a","avatarUrl":"/avatars/080655b2497c2fa44a33da3ccfabb84b.svg","isPro":false,"fullname":"Jiaming Zhou","user":"Jiaming2472","type":"user"},{"_id":"64658e94611ae99d14d5b28d","avatarUrl":"/avatars/1a0d125e9fa6a836e29d7bdb147e9806.svg","isPro":false,"fullname":"Conner Qiu","user":"CQ097","type":"user"},{"_id":"69a396c634ec7f83af77249f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69a396c634ec7f83af77249f/mWFHLavX4_yvGnAU8q-og.jpeg","isPro":false,"fullname":"Yuzhuo Ao","user":"Supramundaner","type":"user"},{"_id":"65d0d5205de2b19b9f40ef46","avatarUrl":"/avatars/3d2609afd19f8c6ac5ac12f523e425af.svg","isPro":false,"fullname":"Cunxin Fan","user":"alfayoung","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"687363d49a81c7dcbcfa2d84","avatarUrl":"/avatars/5d943a5c811ed931c3fdcfee19253049.svg","isPro":false,"fullname":"jj","user":"realman123","type":"user"},{"_id":"65d9be67be18bfea69c63830","avatarUrl":"/avatars/fe68775d214b76f8812db0d066d5be63.svg","isPro":false,"fullname":"Jialong Sun","user":"Pillow-1","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a8eb8c2b938953de99e7939","name":"Robbyant-Research","fullname":"Robbyant Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/sGSUfQbY1ezHjHO7mgSXL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26103.md","query":{}}">
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Abstract
Zero-WAM enables robotic manipulation of unseen tasks by conditioning a causal video-action model on in-context human video guidance, supported by an automatically generated dataset and a future-chunk prediction objective.
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.26103 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.26103 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.26103 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.