Hugging Face Daily Papers · · 3 min read

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<video src=\"https://cdn-uploads.huggingface.co/production/uploads/64636b2551fa6e6306046293/pazYn-AP4rhyzDZ1gAJM3.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>","updatedAt":"2026-08-31T06:43:41.046Z","author":{"_id":"64636b2551fa6e6306046293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64636b2551fa6e6306046293/Uuz6z2MZb_LKLGM8uxF9s.jpeg","fullname":"Hanoona Rasheed","name":"Hanoona","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5921814441680908},"editors":["Hanoona"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64636b2551fa6e6306046293/Uuz6z2MZb_LKLGM8uxF9s.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.28192","authors":[{"_id":"6a9521a2073195fee51572e5","name":"Hanoona Rasheed","hidden":false},{"_id":"6a9521a2073195fee51572e6","name":"Haania Siddiqui","hidden":false},{"_id":"6a9521a2073195fee51572e7","name":"Ming-Hsuan Yang","hidden":false},{"_id":"6a9521a2073195fee51572e8","name":"Fahad Shahbaz Khan","hidden":false},{"_id":"6a9521a2073195fee51572e9","name":"Salman Khan","hidden":false}],"publishedAt":"2026-08-28T00:00:00.000Z","submittedOnDailyAt":"2026-08-31T00:00:00.000Z","title":"Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding","submittedOnDailyBy":{"_id":"64636b2551fa6e6306046293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64636b2551fa6e6306046293/Uuz6z2MZb_LKLGM8uxF9s.jpeg","isPro":false,"fullname":"Hanoona Rasheed","user":"Hanoona","type":"user","name":"Hanoona"},"summary":"Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajectory-level dependencies, reducing the sequential decoding depth to a fixed 1 + 1 rounds, independent of tube length. To enable parallel spatial generation, we introduce Decoupled Block Attention, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry. On VidSTG, PTD reduces Tube Completion Latency by 79x and increases spatial decoding throughput by 92x over standard autoregressive decoding, while also improving grounding accuracy. With a compact 4B backbone, our model performs favorably well on VidSTG and HC-STVG, and generalizes zero-shot to temporal grounding, grounded VideoQA, and referring video object tracking. Our results show parallel tube generation is an efficient and effective alternative to autoregressive localization in videos.","upvotes":8,"discussionId":"6a9521a2073195fee51572ea","projectPage":"https://mbzuai-oryx.github.io/ParallelTubeDecoding/","githubRepo":"https://github.com/mbzuai-oryx/ParallelTubeDecoding","githubRepoAddedBy":"user","ai_summary":"Parallel Tube Decoding enables simultaneous spatial and temporal video grounding by removing autoregressive dependencies, drastically cutting latency while improving accuracy.","ai_keywords":["Spatio-temporal video grounding","Parallel Tube Decoding","Decoupled Block Attention","localization-aware policy optimization","autoregressive decoding","tube completion latency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"61fb9e24dc607a42af5f193f","name":"MBZUAI","fullname":"Mohamed Bin Zayed University of Artificial Intelligence","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1643879908583-603ab5664a944b99e81476e8.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64636b2551fa6e6306046293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64636b2551fa6e6306046293/Uuz6z2MZb_LKLGM8uxF9s.jpeg","isPro":false,"fullname":"Hanoona Rasheed","user":"Hanoona","type":"user"},{"_id":"64807585856901b0edb8d68b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/kmCo6brXhQ7SyxVx3lpsx.jpeg","isPro":false,"fullname":"Muhammad Maaz","user":"mmaaz60","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"68808e1790413512e4f74d25","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68808e1790413512e4f74d25/ALeGZOrYQhOWuuEaslv9i.png","isPro":false,"fullname":"Ankan Deria","user":"ankanmbz","type":"user"},{"_id":"6465e25763e7e09dd02eeede","avatarUrl":"/avatars/3c77d6044ab6db1d29a5ce860a918d4b.svg","isPro":true,"fullname":"Chao Qin","user":"AlfredQin","type":"user"},{"_id":"62440ccec661f366527d05ed","avatarUrl":"/avatars/851c61780b703b02d86651205989723c.svg","isPro":false,"fullname":"Salman Khan","user":"salmaneme","type":"user"},{"_id":"66a788dc34ae1d4c78c8cdbb","avatarUrl":"/avatars/01aeeff26873c4cfd977e8be1b4b04de.svg","isPro":false,"fullname":"Roba Al Majzoub","user":"musk007","type":"user"},{"_id":"6407e5294edf9f5c4fd32228","avatarUrl":"/avatars/8e2d55460e9fe9c426eb552baf4b2cb0.svg","isPro":false,"fullname":"Stoney Kang","user":"sikang99","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61fb9e24dc607a42af5f193f","name":"MBZUAI","fullname":"Mohamed Bin Zayed University of Artificial Intelligence","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1643879908583-603ab5664a944b99e81476e8.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.28192.md","query":{}}">
Papers
arxiv:2608.28192

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Published on Aug 28
· Submitted by
Hanoona Rasheed
on Aug 31
Authors:
,

Abstract

Parallel Tube Decoding enables simultaneous spatial and temporal video grounding by removing autoregressive dependencies, drastically cutting latency while improving accuracy.

Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajectory-level dependencies, reducing the sequential decoding depth to a fixed 1 + 1 rounds, independent of tube length. To enable parallel spatial generation, we introduce Decoupled Block Attention, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry. On VidSTG, PTD reduces Tube Completion Latency by 79x and increases spatial decoding throughput by 92x over standard autoregressive decoding, while also improving grounding accuracy. With a compact 4B backbone, our model performs favorably well on VidSTG and HC-STVG, and generalizes zero-shot to temporal grounding, grounded VideoQA, and referring video object tracking. Our results show parallel tube generation is an efficient and effective alternative to autoregressive localization in videos.

Community

Paper submitter about 1 hour ago

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.28192
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.28192 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.28192 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.28192 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers