Hugging Face Daily Papers · · 3 min read

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Speech Animation using Video Diffusion Model </p>\n<p>Project page : {<a href=\"https://serin-yoon.github.io/projects/anytalk/%7D\" rel=\"nofollow\">https://serin-yoon.github.io/projects/anytalk/}</a><br>Code : {<a href=\"https://github.com/kwanyun/AnyTalk_CsF%7D\" rel=\"nofollow\">https://github.com/kwanyun/AnyTalk_CsF}</a></p>\n","updatedAt":"2026-08-18T04:32:11.210Z","author":{"_id":"639d445524af4747d8d2af52","avatarUrl":"/avatars/6ca870edc993fd3db6b0db4e4848ef7a.svg","fullname":"kwan yun","name":"kwanY","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.562128484249115},"editors":["kwanY"],"editorAvatarUrls":["/avatars/6ca870edc993fd3db6b0db4e4848ef7a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16143","authors":[{"_id":"6a83df57675db694db8cd5d2","name":"Kwan Yun","hidden":false},{"_id":"6a83df57675db694db8cd5d3","name":"Serin Yoon","hidden":false},{"_id":"6a83df57675db694db8cd5d4","name":"Sunjin Jung","hidden":false},{"_id":"6a83df57675db694db8cd5d5","name":"Jung Eun Yoo","hidden":false},{"_id":"6a83df57675db694db8cd5d6","name":"Inyup Lee","hidden":false},{"_id":"6a83df57675db694db8cd5d7","name":"Junyong Noh","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/639d445524af4747d8d2af52/xDkBdT2iYv1bJ11gihg55.png"],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model","submittedOnDailyBy":{"_id":"639d445524af4747d8d2af52","avatarUrl":"/avatars/6ca870edc993fd3db6b0db4e4848ef7a.svg","isPro":false,"fullname":"kwan yun","user":"kwanY","type":"user","name":"kwanY"},"summary":"We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (CsF) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing \"no motion\"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, AnyTalk_{RT}, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.","upvotes":3,"discussionId":"6a83df58675db694db8cd5d8","projectPage":"https://serin-yoon.github.io/projects/anytalk/","githubRepo":"https://github.com/kwanyun/AnyTalk_CsF","githubRepoAddedBy":"user","ai_summary":"AnyTalk generates 3D speech animations for arbitrary characters without animation data by adapting video diffusion models via character-specific fine-tuning and optimizing blendshape parameters from synthesized talking-head videos, with a distilled real-time variant.","ai_keywords":["video diffusion models","Character-specific Fine-tuning (CsF)","blendshape parameters","optimization process","AnyTalk_RT","real-time performance","talking-head video generation","audio-driven speech animation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"635b304962fb2bc1b52c6291","name":"KAIST","fullname":"KAIST","avatar":"https://www.gravatar.com/avatar/eba12517b1eaa0552a14abf582540dbd?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"639d445524af4747d8d2af52","avatarUrl":"/avatars/6ca870edc993fd3db6b0db4e4848ef7a.svg","isPro":false,"fullname":"kwan yun","user":"kwanY","type":"user"},{"_id":"69eb497358d168595f63c670","avatarUrl":"/avatars/14711474a055b9cb670381e2a8c0be63.svg","isPro":false,"fullname":"Junhyuk Jeon","user":"jnhykjeon","type":"user"},{"_id":"67f47de8e34ae64c2108474e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mjGDTtwrgHblkXTKAU6YS.png","isPro":false,"fullname":"Hyunwoo Kwak","user":"Nwoo0413","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"635b304962fb2bc1b52c6291","name":"KAIST","fullname":"KAIST","avatar":"https://www.gravatar.com/avatar/eba12517b1eaa0552a14abf582540dbd?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16143.md","query":{}}">
Papers
arxiv:2608.16143

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

Published on Aug 17
· Submitted by
kwan yun
on Aug 18
Authors:
,

Abstract

AnyTalk generates 3D speech animations for arbitrary characters without animation data by adapting video diffusion models via character-specific fine-tuning and optimizing blendshape parameters from synthesized talking-head videos, with a distilled real-time variant.

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (CsF) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, AnyTalk_{RT}, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.

Community

Paper submitter about 4 hours ago

Speech Animation using Video Diffusion Model

Project page : {https://serin-yoon.github.io/projects/anytalk/}
Code : {https://github.com/kwanyun/AnyTalk_CsF}

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16143
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16143 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16143 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16143 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers