Hugging Face Daily Papers · · 4 min read

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Code: <a href=\"https://github.com/SII-YuanyangYin/Evoke\" rel=\"nofollow\">https://github.com/SII-YuanyangYin/Evoke</a><br>Page: <a href=\"https://evoke-world.github.io/Evoke/\" rel=\"nofollow\">https://evoke-world.github.io/Evoke/</a><br>YouTube Demo: <a href=\"https://youtu.be/QX7PBBaBGdc\" rel=\"nofollow\">https://youtu.be/QX7PBBaBGdc</a></p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/659cb2671d398a2381625b2f/WE9czRRM_RxDx-Kx8ZEmQ.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-08-14T05:39:18.978Z","author":{"_id":"659cb2671d398a2381625b2f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/659cb2671d398a2381625b2f/-y_odTNSgvlcABJ-9GeFf.jpeg","fullname":"SII-YuanyangYin","name":"SII-YuanyangYin","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.38173913955688477},"editors":["SII-YuanyangYin"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/659cb2671d398a2381625b2f/-y_odTNSgvlcABJ-9GeFf.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.13546","authors":[{"_id":"6a7e786542823931a1f17629","name":"Yuanyang Yin","hidden":false},{"_id":"6a7e786542823931a1f1762a","name":"Gongxuan Wang","hidden":false},{"_id":"6a7e786542823931a1f1762b","name":"Yifan Zhan","hidden":false},{"_id":"6a7e786542823931a1f1762c","name":"Chuanhao Li","hidden":false},{"_id":"6a7e786542823931a1f1762d","name":"Kaipeng Zhang","hidden":false},{"_id":"6a7e786542823931a1f1762e","name":"Feng Zhao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/fROWJc_NowCClITP9Y7Mx.mp4"],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"Alaya-EVOKE: From Linear-Scaling Supervision to Endless World","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at 384times 640, each 1.5,s chunk is generated in 2.11,s. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.","upvotes":71,"discussionId":"6a7e786542823931a1f1762f","projectPage":"https://evoke-world.github.io/Evoke/","ai_summary":"Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency.","ai_keywords":["world models","denoiser context","key-value cache","external world state bank","sparse attention","chunk-wise grouping","linear attention","self-forced rollouts","classifier-free guidance","distribution-matching objective"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"659cb2671d398a2381625b2f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/659cb2671d398a2381625b2f/-y_odTNSgvlcABJ-9GeFf.jpeg","isPro":false,"fullname":"SII-YuanyangYin","user":"SII-YuanyangYin","type":"user"},{"_id":"66e04ea2d64de03c6dcb7c29","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66e04ea2d64de03c6dcb7c29/YNF4A2adZAimcyYcoVZg8.jpeg","isPro":false,"fullname":"SII-GongxuanWang","user":"logan1123","type":"user"},{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","isPro":false,"fullname":"Chuanhao","user":"ChuanhaoLi","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"63468720dd6d90d82ccf3450","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63468720dd6d90d82ccf3450/tVBFlmZNz8FRMkOrDaDID.jpeg","isPro":false,"fullname":"YSH","user":"BestWishYsh","type":"user"},{"_id":"6a7c2986ef515a0aa78152ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a7c2986ef515a0aa78152ff/hPFNrsFALJtDvsX7GeFAB.jpeg","isPro":false,"fullname":"Liam Doyle","user":"danielcybq","type":"user"},{"_id":"6a7d1fcb98009700bd8f3387","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a7d1fcb98009700bd8f3387/0BLfUSUgAxeNJ-9bb3Jva.jpeg","isPro":false,"fullname":"井上 結衣","user":"sakurakobayashi","type":"user"},{"_id":"6a7d26a5813097f36c653f81","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a7d26a5813097f36c653f81/2SXDCSMYi81B8tZamYEsB.jpeg","isPro":false,"fullname":"Hazel Wilson","user":"joshuagonzalez","type":"user"},{"_id":"65f1713552c38a91e0a445e8","avatarUrl":"/avatars/47ab3ada51c9b9976ac1cd0c4301c373.svg","isPro":false,"fullname":"kaipeng","user":"kpzhang996","type":"user"},{"_id":"63342778d92c5842ae728aef","avatarUrl":"/avatars/888eb265643633c5fdd7048be9bfe98f.svg","isPro":false,"fullname":"Fengbo Lan","user":"fblan","type":"user"},{"_id":"676bc71e490f3664721e81eb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/DRAdleyZ_Zy5h4IldZ3gb.png","isPro":false,"fullname":"Sakura Sato","user":"SakuraSato","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.13546.md","query":{}}">
Papers
arxiv:2608.13546

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Published on Aug 13
· Submitted by
taesiri
on Aug 14
#1 Paper of the day
Authors:
,

Abstract

Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency.

Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at 384times 640, each 1.5,s chunk is generated in 2.11,s. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.13546
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.13546 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.13546 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.13546 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers