<strong>What if VLA action decoders knew not only what action to execute, but also what the behavior is trying to achieve?</strong><br>We introduce Intention Distillation (INDI), which distills behavior-level intent from a frozen teacher VLM into the VLA action decoder, while requiring no teacher at deployment. INDI consistently improves GR00T-N1.7 and π₀.₅ across simulation and real-world manipulation, including 64.3% → 84.7% on SimplerEnv-Bridge and gains of up to +12.0 pp on longer-horizon real-world tasks.</p>\n","updatedAt":"2026-08-31T11:45:54.217Z","author":{"_id":"6703b69c4a471e82db829c7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg","fullname":"Sangoh Lee","name":"leesangoh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9214158058166504},"editors":["leesangoh"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23478","authors":[{"_id":"6a9504da073195fee515728e","user":{"_id":"6703b69c4a471e82db829c7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg","isPro":false,"fullname":"Sangoh Lee","user":"leesangoh","type":"user","name":"leesangoh"},"name":"Sangoh Lee","status":"claimed_verified","statusLastChangedAt":"2026-08-31T08:38:27.277Z","hidden":false},{"_id":"6a9504da073195fee515728f","name":"Sangwoo Mo","hidden":false},{"_id":"6a9504da073195fee5157290","name":"Wook-Shin Han","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/EMVjU6UDgYB2iDA62yuH6.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/dU6YBBgk17cYHdzrMDRpI.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/BxW-FO3Ol-yCl-vGib6W0.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/OzKwGyyQq1m-uSREmSNqn.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/TI0Gts84qwzFliMfSn9_L.png","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/xpb1wDDSeeS84YFddb9NS.png"],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-31T00:00:00.000Z","title":"Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models","submittedOnDailyBy":{"_id":"6703b69c4a471e82db829c7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg","isPro":false,"fullname":"Sangoh Lee","user":"leesangoh","type":"user","name":"leesangoh"},"summary":"Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on π_{0.5} across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.","upvotes":21,"discussionId":"6a9504da073195fee5157291","projectPage":"https://leesangoh.github.io/indi-project-page/","githubRepo":"https://github.com/Leesangoh/INDI","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"62459012e1b9dab15a3e6674","name":"POSTECH","fullname":"Pohang University of Science and Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648726022705-62458f43d5895bdf34ee7d56.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"648804a275e78c1ebb7755bf","avatarUrl":"/avatars/d7f4fb2c07a4fc2dd72c5aa0b77d08da.svg","isPro":false,"fullname":"Sungho Park","user":"pshlego","type":"user"},{"_id":"681605320144162c256950ae","avatarUrl":"/avatars/c6eadc7aa634fdfeb43f07b241815f7b.svg","isPro":false,"fullname":"kim","user":"jueunkim","type":"user"},{"_id":"678f172a02a1c58b5b3086f6","avatarUrl":"/avatars/ed72e9aacbf75f73406317d0adc00771.svg","isPro":false,"fullname":"Sunho Cha","user":"carprefer","type":"user"},{"_id":"6826afcf08f7cb26de069b5e","avatarUrl":"/avatars/6e0b357df51e516a674edb66889fb86e.svg","isPro":false,"fullname":"Seungah Jang","user":"sajang928","type":"user"},{"_id":"6858f58b48fab0227165bca3","avatarUrl":"/avatars/ac0e7a3d2c225f50c16f7bca68a8ca82.svg","isPro":false,"fullname":"HyoJeong Yun","user":"hyojeongyun","type":"user"},{"_id":"6682138dc7d9ff092299098d","avatarUrl":"/avatars/d160376e55115217d235e2f3353c9e7e.svg","isPro":false,"fullname":"Suchan Lee","user":"isuchan0212","type":"user"},{"_id":"6a956bd0df9edc286650606c","avatarUrl":"/avatars/6a222072ee0ef2a6b7337bb675487586.svg","isPro":false,"fullname":"Wonseok Lee","user":"wslee2265","type":"user"},{"_id":"64ad53921db46d71a5c576e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ad53921db46d71a5c576e5/mAPg4l6kZ6x1dvoxdPRm3.png","isPro":true,"fullname":"ChuGyouk","user":"ChuGyouk","type":"user"},{"_id":"6a956cc6f60142cfe08663cd","avatarUrl":"/avatars/c39389970c02220be61abe17ba41dee6.svg","isPro":false,"fullname":"Junho Moon","user":"JunhoMoon0","type":"user"},{"_id":"660bf2e99069ffa7baec28da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660bf2e99069ffa7baec28da/XAstSmX2b7UPB6Mgjm6yY.png","isPro":false,"fullname":"Junyeong Song","user":"junyeong-nero","type":"user"},{"_id":"66b0d9f38c0bf54236f22603","avatarUrl":"/avatars/6c7f04397f6f09d8626e0d714ef3c866.svg","isPro":false,"fullname":"SungWoo Kwon","user":"ksw6895","type":"user"},{"_id":"69b8010bee9371902f047f75","avatarUrl":"/avatars/fdad957b2812372d8391a44183a4035a.svg","isPro":false,"fullname":"Taejun Yoon","user":"tjyoon","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"62459012e1b9dab15a3e6674","name":"POSTECH","fullname":"Pohang University of Science and Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648726022705-62458f43d5895bdf34ee7d56.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23478.md","query":{}}">
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on π_{0.5} across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.
Community
What if VLA action decoders knew not only what action to execute, but also what the behavior is trying to achieve?
We introduce Intention Distillation (INDI), which distills behavior-level intent from a frozen teacher VLM into the VLA action decoder, while requiring no teacher at deployment. INDI consistently improves GR00T-N1.7 and π₀.₅ across simulation and real-world manipulation, including 64.3% → 84.7% on SimplerEnv-Bridge and gains of up to +12.0 pp on longer-horizon real-world tasks.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.23478 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.23478 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.23478 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.