Hugging Face Daily Papers · · 4 min read

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds</p>\n","updatedAt":"2026-08-27T05:38:43.271Z","author":{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","fullname":"Xue Zeyue","name":"xzyhku","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5209053754806519},"editors":["xzyhku"],"editorAvatarUrls":["/avatars/593c14c907848da7dbc9e5418751bd94.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23383","authors":[{"_id":"6a8d0b9f5add2537c32e96f3","name":"Nan Duan","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f4","name":"Haoyang Huang","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f5","user":{"_id":"66608add236f958513d21d2e","avatarUrl":"/avatars/53eca0891c98cbb93be899885160a983.svg","isPro":false,"fullname":"Weiyang Jin","user":"Wayne-King","type":"user","name":"Wayne-King"},"name":"Weiyang Jin","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.853Z","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f6","user":{"_id":"683d173214785f6d5902f9c0","avatarUrl":"/avatars/2357d97bbcecd3a51279442bd99a0c50.svg","isPro":false,"fullname":"Haoran Li","user":"jahnsonblack","type":"user","name":"jahnsonblack"},"name":"Haoran Li","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:53.045Z","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f7","name":"Yaowei Li","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f8","name":"Yuming Li","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f9","name":"Yijun Liu","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fa","name":"Xin Lu","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fb","name":"Xiaoxiao Ma","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fc","name":"Yanwen Ma","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fd","name":"Yaofeng Su","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fe","name":"Yilang Sun","hidden":false},{"_id":"6a8d0b9f5add2537c32e96ff","name":"Haoyu Wang","hidden":false},{"_id":"6a8d0b9f5add2537c32e9700","name":"Zeyue Xue","hidden":false},{"_id":"6a8d0b9f5add2537c32e9701","name":"Songchun Zhang","hidden":false},{"_id":"6a8d0b9f5add2537c32e9702","user":{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user","name":"JunhaoZhuang"},"name":"Junhao Zhuang","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:54.904Z","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds","submittedOnDailyBy":{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","isPro":false,"fullname":"Xue Zeyue","user":"xzyhku","type":"user","name":"xzyhku"},"summary":"Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.","upvotes":19,"discussionId":"6a8d0b9f5add2537c32e9703","projectPage":"https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/","githubRepo":"https://github.com/jd-opensource/JoyAI-Echo","githubRepoAddedBy":"user","ai_summary":"JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences.","ai_keywords":["cross-shot memory","speaker cues","speech-filtered audio","6-DoF camera trajectories","geometry-aware conditioning","bidirectional audio-visual backbone","causal few-step generator","progressive teacher forcing","Self-Gradient Forcing","self-generated rollouts"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1951},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user"},{"_id":"6362801380c1a705a6ea54ac","avatarUrl":"/avatars/041ad5abf9be42e336938f51ebb8746c.svg","isPro":false,"fullname":"Yaowei Li","user":"Yw22","type":"user"},{"_id":"661639c4dd9284ce5d361f0f","avatarUrl":"/avatars/8919a4662abe4198e91c2b5501d2de41.svg","isPro":false,"fullname":"Lirui Zhao","user":"LiruiZhao","type":"user"},{"_id":"683d173214785f6d5902f9c0","avatarUrl":"/avatars/2357d97bbcecd3a51279442bd99a0c50.svg","isPro":false,"fullname":"Haoran Li","user":"jahnsonblack","type":"user"},{"_id":"662f43dbe93bb73804fa1606","avatarUrl":"/avatars/1a3c117283598f566fc59b02e0df2f9e.svg","isPro":true,"fullname":"Ruofeng Yang","user":"RuofengYang","type":"user"},{"_id":"69f71609c0b462260118b120","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/kPUU7CbY1wpFDPVZKSL3E.png","isPro":false,"fullname":"Yonghui Wang","user":"harryyhwang","type":"user"},{"_id":"65781534e390cfd40998d7af","avatarUrl":"/avatars/85d3be7f74b9d959ba9b3ccc04398536.svg","isPro":false,"fullname":"Hongyang Wei","user":"nonwhy","type":"user"},{"_id":"668df98de9e585e8718f767f","avatarUrl":"/avatars/2be52f4ae88a0991c8ae584f8e870734.svg","isPro":false,"fullname":"Xiangyang Luo","user":"XiangyangLuo02","type":"user"},{"_id":"6411c801e872ae3fb1e2c96e","avatarUrl":"/avatars/f8898dc13d700e545eedbbfab1c18353.svg","isPro":true,"fullname":"Franklin","user":"Franklinzhang","type":"user"},{"_id":"646eac510867c99c2d3fde08","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646eac510867c99c2d3fde08/fvIiW7zj4aTNbp16kTNBA.jpeg","isPro":false,"fullname":"Yaofeng Su","user":"Exploration","type":"user"},{"_id":"66608add236f958513d21d2e","avatarUrl":"/avatars/53eca0891c98cbb93be899885160a983.svg","isPro":false,"fullname":"Weiyang Jin","user":"Wayne-King","type":"user"},{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","isPro":false,"fullname":"Xue Zeyue","user":"xzyhku","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23383.md","query":{}}">
Papers
arxiv:2608.23383

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Published on Aug 24
· Submitted by
Xue Zeyue
on Aug 27

Abstract

JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences.

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.

Community

Paper submitter about 3 hours ago

JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23383
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.23383 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.23383 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23383 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers