Hugging Face Daily Papers · · 5 min read

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.</p>\n","updatedAt":"2026-08-27T03:45:25.025Z","author":{"_id":"675be14c40726aac79aa4b87","avatarUrl":"/avatars/0d15b1b46110750a2975b89d9a81e517.svg","fullname":"ZHE LIU","name":"happinessqq","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8258444666862488},"editors":["happinessqq"],"editorAvatarUrls":["/avatars/0d15b1b46110750a2975b89d9a81e517.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26067","authors":[{"_id":"6a8fab862c24e8c5fab32996","name":"Zhe Liu","hidden":false},{"_id":"6a8fab862c24e8c5fab32997","name":"Jinghua Hou","hidden":false},{"_id":"6a8fab862c24e8c5fab32998","user":{"_id":"64b8faeb8b53fb5dbdfecae5","avatarUrl":"/avatars/9f5919600ee69c38be896dd959bb8724.svg","isPro":false,"fullname":"Yuxiang Lu","user":"yxlu0","type":"user","name":"yxlu0"},"name":"Yuxiang Lu","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.959Z","hidden":false},{"_id":"6a8fab862c24e8c5fab32999","name":"Zhenya Yang","hidden":false},{"_id":"6a8fab862c24e8c5fab3299a","name":"Xianzhe Fan","hidden":false},{"_id":"6a8fab862c24e8c5fab3299b","name":"Junwei Luo","hidden":false},{"_id":"6a8fab862c24e8c5fab3299c","name":"Junyi Li","hidden":false},{"_id":"6a8fab862c24e8c5fab3299d","name":"Ruihua Han","hidden":false},{"_id":"6a8fab862c24e8c5fab3299e","name":"Zhi Hou","hidden":false},{"_id":"6a8fab862c24e8c5fab3299f","name":"Hengshuang Zhao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/675be14c40726aac79aa4b87/EqzF7I6TmkyIP5eDgl0Bv.mp4"],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models","submittedOnDailyBy":{"_id":"675be14c40726aac79aa4b87","avatarUrl":"/avatars/0d15b1b46110750a2975b89d9a81e517.svg","isPro":false,"fullname":"ZHE LIU","user":"happinessqq","type":"user","name":"happinessqq"},"summary":"Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.","upvotes":14,"discussionId":"6a8fab862c24e8c5fab329a0","projectPage":"https://happinesslz.github.io/projects/StreamPI/","githubRepo":"https://github.com/hku-sail/StreamPI","githubRepoAddedBy":"user","ai_summary":"StreamPI enhances single-frame vision-language-action models with streaming temporal reasoning via instruction-anchored attention and randomized interval training, improving robot manipulation without extra parameters.","ai_keywords":["Vision-Language-Action models","streaming multimodal temporal modeling","instruction-anchored temporal modeling","bidirectional attention","causal attention","autoregressive streaming inference","random-interval streaming training","length extrapolation","LLM backbone"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"675be14c40726aac79aa4b87","avatarUrl":"/avatars/0d15b1b46110750a2975b89d9a81e517.svg","isPro":false,"fullname":"ZHE LIU","user":"happinessqq","type":"user"},{"_id":"69a14e7ca4617ee699bca39c","avatarUrl":"/avatars/3ce05b41e0c88afbea724520fa1cfc79.svg","isPro":false,"fullname":"Jinghua Hou","user":"almoonysl","type":"user"},{"_id":"64b8faeb8b53fb5dbdfecae5","avatarUrl":"/avatars/9f5919600ee69c38be896dd959bb8724.svg","isPro":false,"fullname":"Yuxiang Lu","user":"yxlu0","type":"user"},{"_id":"640051b8f4ff62c2616ae97d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/640051b8f4ff62c2616ae97d/RcDyY7hN4VuY_3a0Bua6b.jpeg","isPro":false,"fullname":"XianzheFan","user":"XianzheFan","type":"user"},{"_id":"67d7dc2c2acfd67127fbfcf7","avatarUrl":"/avatars/f64d709fc1985bccced7110e7d79c742.svg","isPro":false,"fullname":"Liu Yantong","user":"LittleLisa","type":"user"},{"_id":"697ca0dfc00f332cf466db32","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/_AN9lrPR0plnrUU5-P0jU.jpeg","isPro":false,"fullname":"Yinghao Xiang","user":"MessianX","type":"user"},{"_id":"64e9c855233101ed99ca2315","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/ozwIxW0ofJMR6kj623p1T.png","isPro":false,"fullname":"YANG, Zhenya","user":"ANIYA673","type":"user"},{"_id":"63aaf2a2a4bdd629b7eb2b5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63aaf2a2a4bdd629b7eb2b5b/WOa3nAUNy5D3MsFUV9B8Z.jpeg","isPro":false,"fullname":"Junyi Li","user":"ProvenceStar","type":"user"},{"_id":"686bcb162801c0c29ea1da31","avatarUrl":"/avatars/1c85a2c4ddf5b2bfcbb5a0ffb875a0ac.svg","isPro":false,"fullname":"Xueli Sun","user":"shirley430316","type":"user"},{"_id":"6451364c4b84eee0deea68fa","avatarUrl":"/avatars/cc5c52f9b30f43b7faa47d0d5848a492.svg","isPro":false,"fullname":"Xirui Li","user":"lixirui142","type":"user"},{"_id":"69bcb9b54b067234c61b1fe0","avatarUrl":"/avatars/0e4f620eeda2f9deec344a4cf1524ba5.svg","isPro":false,"fullname":"Ruihua Han","user":"hanrobot","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26067.md","query":{}}">
Papers
arxiv:2608.26067

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Published on Aug 26
· Submitted by
ZHE LIU
on Aug 27
Authors:
,

Abstract

StreamPI enhances single-frame vision-language-action models with streaming temporal reasoning via instruction-anchored attention and randomized interval training, improving robot manipulation without extra parameters.

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.

Community

Paper submitter about 5 hours ago

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.26067
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.26067 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.26067 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.26067 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers