JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds</p>\n","updatedAt":"2026-08-27T05:38:43.271Z","author":{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","fullname":"Xue Zeyue","name":"xzyhku","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5209053754806519},"editors":["xzyhku"],"editorAvatarUrls":["/avatars/593c14c907848da7dbc9e5418751bd94.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23383","authors":[{"_id":"6a8d0b9f5add2537c32e96f3","name":"Nan Duan","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f4","name":"Haoyang Huang","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f5","user":{"_id":"66608add236f958513d21d2e","avatarUrl":"/avatars/53eca0891c98cbb93be899885160a983.svg","isPro":false,"fullname":"Weiyang Jin","user":"Wayne-King","type":"user","name":"Wayne-King"},"name":"Weiyang Jin","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.853Z","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f6","user":{"_id":"683d173214785f6d5902f9c0","avatarUrl":"/avatars/2357d97bbcecd3a51279442bd99a0c50.svg","isPro":false,"fullname":"Haoran Li","user":"jahnsonblack","type":"user","name":"jahnsonblack"},"name":"Haoran Li","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:53.045Z","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f7","name":"Yaowei Li","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f8","name":"Yuming Li","hidden":false},{"_id":"6a8d0b9f5add2537c32e96f9","name":"Yijun Liu","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fa","name":"Xin Lu","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fb","name":"Xiaoxiao Ma","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fc","name":"Yanwen Ma","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fd","name":"Yaofeng Su","hidden":false},{"_id":"6a8d0b9f5add2537c32e96fe","name":"Yilang Sun","hidden":false},{"_id":"6a8d0b9f5add2537c32e96ff","name":"Haoyu Wang","hidden":false},{"_id":"6a8d0b9f5add2537c32e9700","name":"Zeyue Xue","hidden":false},{"_id":"6a8d0b9f5add2537c32e9701","name":"Songchun Zhang","hidden":false},{"_id":"6a8d0b9f5add2537c32e9702","user":{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user","name":"JunhaoZhuang"},"name":"Junhao Zhuang","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:54.904Z","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds","submittedOnDailyBy":{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","isPro":false,"fullname":"Xue Zeyue","user":"xzyhku","type":"user","name":"xzyhku"},"summary":"Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.","upvotes":19,"discussionId":"6a8d0b9f5add2537c32e9703","projectPage":"https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/","githubRepo":"https://github.com/jd-opensource/JoyAI-Echo","githubRepoAddedBy":"user","ai_summary":"JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences.","ai_keywords":["cross-shot memory","speaker cues","speech-filtered audio","6-DoF camera trajectories","geometry-aware conditioning","bidirectional audio-visual backbone","causal few-step generator","progressive teacher forcing","Self-Gradient Forcing","self-generated rollouts"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1951},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user"},{"_id":"6362801380c1a705a6ea54ac","avatarUrl":"/avatars/041ad5abf9be42e336938f51ebb8746c.svg","isPro":false,"fullname":"Yaowei Li","user":"Yw22","type":"user"},{"_id":"661639c4dd9284ce5d361f0f","avatarUrl":"/avatars/8919a4662abe4198e91c2b5501d2de41.svg","isPro":false,"fullname":"Lirui Zhao","user":"LiruiZhao","type":"user"},{"_id":"683d173214785f6d5902f9c0","avatarUrl":"/avatars/2357d97bbcecd3a51279442bd99a0c50.svg","isPro":false,"fullname":"Haoran Li","user":"jahnsonblack","type":"user"},{"_id":"662f43dbe93bb73804fa1606","avatarUrl":"/avatars/1a3c117283598f566fc59b02e0df2f9e.svg","isPro":true,"fullname":"Ruofeng Yang","user":"RuofengYang","type":"user"},{"_id":"69f71609c0b462260118b120","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/kPUU7CbY1wpFDPVZKSL3E.png","isPro":false,"fullname":"Yonghui Wang","user":"harryyhwang","type":"user"},{"_id":"65781534e390cfd40998d7af","avatarUrl":"/avatars/85d3be7f74b9d959ba9b3ccc04398536.svg","isPro":false,"fullname":"Hongyang Wei","user":"nonwhy","type":"user"},{"_id":"668df98de9e585e8718f767f","avatarUrl":"/avatars/2be52f4ae88a0991c8ae584f8e870734.svg","isPro":false,"fullname":"Xiangyang Luo","user":"XiangyangLuo02","type":"user"},{"_id":"6411c801e872ae3fb1e2c96e","avatarUrl":"/avatars/f8898dc13d700e545eedbbfab1c18353.svg","isPro":true,"fullname":"Franklin","user":"Franklinzhang","type":"user"},{"_id":"646eac510867c99c2d3fde08","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646eac510867c99c2d3fde08/fvIiW7zj4aTNbp16kTNBA.jpeg","isPro":false,"fullname":"Yaofeng Su","user":"Exploration","type":"user"},{"_id":"66608add236f958513d21d2e","avatarUrl":"/avatars/53eca0891c98cbb93be899885160a983.svg","isPro":false,"fullname":"Weiyang Jin","user":"Wayne-King","type":"user"},{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","isPro":false,"fullname":"Xue Zeyue","user":"xzyhku","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23383.md","query":{}}">
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Abstract
JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences.
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
Community
JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.23383 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.23383 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.23383 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.