Hugging Face Daily Papers · · 4 min read

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<strong>What if VLA action decoders knew not only what action to execute, but also what the behavior is trying to achieve?</strong><br>We introduce Intention Distillation (INDI), which distills behavior-level intent from a frozen teacher VLM into the VLA action decoder, while requiring no teacher at deployment. INDI consistently improves GR00T-N1.7 and π₀.₅ across simulation and real-world manipulation, including 64.3% → 84.7% on SimplerEnv-Bridge and gains of up to +12.0 pp on longer-horizon real-world tasks.</p>\n","updatedAt":"2026-08-31T11:45:54.217Z","author":{"_id":"6703b69c4a471e82db829c7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg","fullname":"Sangoh Lee","name":"leesangoh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9214158058166504},"editors":["leesangoh"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23478","authors":[{"_id":"6a9504da073195fee515728e","user":{"_id":"6703b69c4a471e82db829c7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg","isPro":false,"fullname":"Sangoh Lee","user":"leesangoh","type":"user","name":"leesangoh"},"name":"Sangoh Lee","status":"claimed_verified","statusLastChangedAt":"2026-08-31T08:38:27.277Z","hidden":false},{"_id":"6a9504da073195fee515728f","name":"Sangwoo Mo","hidden":false},{"_id":"6a9504da073195fee5157290","name":"Wook-Shin Han","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/EMVjU6UDgYB2iDA62yuH6.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/dU6YBBgk17cYHdzrMDRpI.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/BxW-FO3Ol-yCl-vGib6W0.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/OzKwGyyQq1m-uSREmSNqn.mp4","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/TI0Gts84qwzFliMfSn9_L.png","https://cdn-uploads.huggingface.co/production/uploads/6703b69c4a471e82db829c7a/xpb1wDDSeeS84YFddb9NS.png"],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-31T00:00:00.000Z","title":"Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models","submittedOnDailyBy":{"_id":"6703b69c4a471e82db829c7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6703b69c4a471e82db829c7a/aXoQ86gpcTNFG-o9Aba1H.jpeg","isPro":false,"fullname":"Sangoh Lee","user":"leesangoh","type":"user","name":"leesangoh"},"summary":"Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on π_{0.5} across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.","upvotes":21,"discussionId":"6a9504da073195fee5157291","projectPage":"https://leesangoh.github.io/indi-project-page/","githubRepo":"https://github.com/Leesangoh/INDI","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"62459012e1b9dab15a3e6674","name":"POSTECH","fullname":"Pohang University of Science and Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648726022705-62458f43d5895bdf34ee7d56.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"648804a275e78c1ebb7755bf","avatarUrl":"/avatars/d7f4fb2c07a4fc2dd72c5aa0b77d08da.svg","isPro":false,"fullname":"Sungho Park","user":"pshlego","type":"user"},{"_id":"681605320144162c256950ae","avatarUrl":"/avatars/c6eadc7aa634fdfeb43f07b241815f7b.svg","isPro":false,"fullname":"kim","user":"jueunkim","type":"user"},{"_id":"678f172a02a1c58b5b3086f6","avatarUrl":"/avatars/ed72e9aacbf75f73406317d0adc00771.svg","isPro":false,"fullname":"Sunho Cha","user":"carprefer","type":"user"},{"_id":"6826afcf08f7cb26de069b5e","avatarUrl":"/avatars/6e0b357df51e516a674edb66889fb86e.svg","isPro":false,"fullname":"Seungah Jang","user":"sajang928","type":"user"},{"_id":"6858f58b48fab0227165bca3","avatarUrl":"/avatars/ac0e7a3d2c225f50c16f7bca68a8ca82.svg","isPro":false,"fullname":"HyoJeong Yun","user":"hyojeongyun","type":"user"},{"_id":"6682138dc7d9ff092299098d","avatarUrl":"/avatars/d160376e55115217d235e2f3353c9e7e.svg","isPro":false,"fullname":"Suchan Lee","user":"isuchan0212","type":"user"},{"_id":"6a956bd0df9edc286650606c","avatarUrl":"/avatars/6a222072ee0ef2a6b7337bb675487586.svg","isPro":false,"fullname":"Wonseok Lee","user":"wslee2265","type":"user"},{"_id":"64ad53921db46d71a5c576e5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ad53921db46d71a5c576e5/mAPg4l6kZ6x1dvoxdPRm3.png","isPro":true,"fullname":"ChuGyouk","user":"ChuGyouk","type":"user"},{"_id":"6a956cc6f60142cfe08663cd","avatarUrl":"/avatars/c39389970c02220be61abe17ba41dee6.svg","isPro":false,"fullname":"Junho Moon","user":"JunhoMoon0","type":"user"},{"_id":"660bf2e99069ffa7baec28da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660bf2e99069ffa7baec28da/XAstSmX2b7UPB6Mgjm6yY.png","isPro":false,"fullname":"Junyeong Song","user":"junyeong-nero","type":"user"},{"_id":"66b0d9f38c0bf54236f22603","avatarUrl":"/avatars/6c7f04397f6f09d8626e0d714ef3c866.svg","isPro":false,"fullname":"SungWoo Kwon","user":"ksw6895","type":"user"},{"_id":"69b8010bee9371902f047f75","avatarUrl":"/avatars/fdad957b2812372d8391a44183a4035a.svg","isPro":false,"fullname":"Taejun Yoon","user":"tjyoon","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"62459012e1b9dab15a3e6674","name":"POSTECH","fullname":"Pohang University of Science and Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1648726022705-62458f43d5895bdf34ee7d56.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23478.md","query":{}}">
Papers
arxiv:2608.23478

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Published on Aug 24
· Submitted by
Sangoh Lee
on Aug 31
Authors:

Abstract

Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on π_{0.5} across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

Community

Paper author Paper submitter about 3 hours ago

What if VLA action decoders knew not only what action to execute, but also what the behavior is trying to achieve?
We introduce Intention Distillation (INDI), which distills behavior-level intent from a frozen teacher VLM into the VLA action decoder, while requiring no teacher at deployment. INDI consistently improves GR00T-N1.7 and π₀.₅ across simulation and real-world manipulation, including 64.3% → 84.7% on SimplerEnv-Bridge and gains of up to +12.0 pp on longer-horizon real-world tasks.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23478
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.23478 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.23478 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23478 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers