Hugging Face Daily Papers · · 8 min read

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀 We are excited to introduce <strong>VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation</strong>, the <strong>first reward model specifically designed for joint video-audio generation</strong>.</p>\n<p>Existing metrics typically evaluate video and audio separately, overlooking the holistic cross-modal coherence that shapes human preference. When used for post-training, these fragmented metrics may also lead to reward hacking, where metric scores improve without a corresponding improvement in perceptual quality.</p>\n<p>Our key contributions include:</p>\n<ul>\n<li><strong>VA-Judger</strong>, a reasoning-based omni-modal reward model that jointly assesses visual and audio quality, text alignment, audio-video synchronization, semantic coherence, and overall human preference.</li>\n<li><strong>VAPref-10K</strong> and <strong>VA-Judger-Bench</strong>, which provide human preference annotations and a challenging benchmark covering both in-domain and out-of-domain video-audio generation models.</li>\n<li>A complete reward-modeling and post-training framework that uses VA-Judger to improve joint video-audio generation.</li>\n</ul>\n<p>VA-Judger substantially outperforms single-dimensional metrics and omni-modal model baselines such as Qwen3-Omni. It also generalizes reliably to unseen closed-source generation models.</p>\n<p>When used to post-train LTX-2, the resulting model achieves a <strong>62.30% human preference rate</strong>, compared with <strong>27.63%</strong> for the OmniNFT-trained version and <strong>10.08%</strong> for the original LTX-2. It also achieves the best performance on <strong>11 out of 13 objective metrics</strong>.</p>\n<p>🔗 Project: <a href=\"https://sharelab-sii.github.io/VA-Judger/\" rel=\"nofollow\">https://sharelab-sii.github.io/VA-Judger/</a><br>📄 Paper: <a href=\"https://arxiv.org/abs/2608.18607\" rel=\"nofollow\">https://arxiv.org/abs/2608.18607</a><br>💻 Code: <a href=\"https://github.com/ShareLab-SII/VA-Judger\" rel=\"nofollow\">https://github.com/ShareLab-SII/VA-Judger</a><br>🤗 Models: <a href=\"https://huggingface.co/ShareLab-SII/VA-Judger\">https://huggingface.co/ShareLab-SII/VA-Judger</a><br>📊 Dataset: <a href=\"https://huggingface.co/datasets/ShareLab-SII/VA-Judger-Bench\">https://huggingface.co/datasets/ShareLab-SII/VA-Judger-Bench</a><br>🎮 Demo: <a href=\"https://www.youtube.com/watch?v=HUiEFLTY9-E\" rel=\"nofollow\">https://www.youtube.com/watch?v=HUiEFLTY9-E</a></p>\n<p>Further training code and the full VAPref-10K dataset will be released soon. Stay tuned!</p>\n","updatedAt":"2026-08-21T00:22:54.633Z","author":{"_id":"6809f1c00046f092a261be78","avatarUrl":"/avatars/4cadf9bdd6e0f8adf39f38eb3aaa9419.svg","fullname":"Yinming Huang","name":"YinmingHuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8244888782501221},"editors":["YinmingHuang"],"editorAvatarUrls":["/avatars/4cadf9bdd6e0f8adf39f38eb3aaa9419.svg"],"reactions":[],"isReport":false}},{"id":"6a87ab2de41d08af65cf4473","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-21T01:34:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward](https://huggingface.co/papers/2608.06930) (2026)\n* [AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning](https://huggingface.co/papers/2608.09559) (2026)\n* [OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation](https://huggingface.co/papers/2607.23855) (2026)\n* [AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning](https://huggingface.co/papers/2607.12820) (2026)\n* [OmniReasoner: Thinking with Long Audio-Video via Native Tool Use](https://huggingface.co/papers/2607.19339) (2026)\n* [Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning](https://huggingface.co/papers/2608.02831) (2026)\n* [DiT-Reward: Generative Representations for Text-to-Image Reward Modeling](https://huggingface.co/papers/2606.23626) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2608.06930\">AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.09559\">AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.23855\">OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.12820\">AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.19339\">OmniReasoner: Thinking with Long Audio-Video via Native Tool Use</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.02831\">Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.23626\">DiT-Reward: Generative Representations for Text-to-Image Reward Modeling</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-21T01:34:37.154Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7358484864234924},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.18607","authors":[{"_id":"6a879a4889e517cbfd75db83","name":"Yinming Huang","hidden":false},{"_id":"6a879a4889e517cbfd75db84","name":"Shuyuan Tu","hidden":false},{"_id":"6a879a4889e517cbfd75db85","name":"Xi Yan","hidden":false},{"_id":"6a879a4889e517cbfd75db86","name":"Zihan Yang","hidden":false},{"_id":"6a879a4889e517cbfd75db87","name":"Jianhua Han","hidden":false},{"_id":"6a879a4889e517cbfd75db88","name":"Xu Hang","hidden":false},{"_id":"6a879a4889e517cbfd75db89","name":"Yu-Gang Jiang","hidden":false},{"_id":"6a879a4889e517cbfd75db8a","name":"Zuxuan Wu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6809f1c00046f092a261be78/ptZfRDj8NkfLFQ7wXfdxq.mp4"],"publishedAt":"2026-08-19T00:00:00.000Z","submittedOnDailyAt":"2026-08-20T00:00:00.000Z","title":"VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation","submittedOnDailyBy":{"_id":"6809f1c00046f092a261be78","avatarUrl":"/avatars/4cadf9bdd6e0f8adf39f38eb3aaa9419.svg","isPro":false,"fullname":"Yinming Huang","user":"YinmingHuang","type":"user","name":"YinmingHuang"},"summary":"Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.","upvotes":5,"discussionId":"6a879a4989e517cbfd75db8b","ai_summary":"A human-aligned chain-of-thought reward model and preference dataset improve joint video-audio generation by replacing fragmented metrics with coherent, dimension-wise reinforcement learning.","ai_keywords":["reinforcement learning","joint video-audio generation","reward hacking","human-preference dataset","VAPref-10K","VA-Judger-Bench","VA-Judger","chain-of-thought","omni-reward model","rejection sampling","dimension-wise reinforcement learning"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"69aebd93ccae08982f36adae","name":"ShareLab-SII","fullname":"ShareLab-SII","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d4b07fa99b113364b9ca86/nTBRWVJ08tHVh-mh6i8tL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6809f1c00046f092a261be78","avatarUrl":"/avatars/4cadf9bdd6e0f8adf39f38eb3aaa9419.svg","isPro":false,"fullname":"Yinming Huang","user":"YinmingHuang","type":"user"},{"_id":"66da6972eae491c64243e8f3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/SgX1j3QGKMk_YFTnM9Wr_.png","isPro":false,"fullname":"Shuyuan Tu","user":"FrancisRing","type":"user"},{"_id":"6a48d4842ecab0cec534acd6","avatarUrl":"/avatars/79b8c880c3c1dcc5fcb0ce610e9c1ca4.svg","isPro":false,"fullname":"kaka","user":"kakawby","type":"user"},{"_id":"64a6a1ecb1f875f03fa05a8d","avatarUrl":"/avatars/4ca06e57d43c66c76b71070f02003d7b.svg","isPro":false,"fullname":"Zihan Yang","user":"ymyy307","type":"user"},{"_id":"6690e13ccbcaf7ab0ec1c971","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/e8KDV6J29tviXlIpLZPq6.png","isPro":false,"fullname":"Tony.Li","user":"lkdhy","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69aebd93ccae08982f36adae","name":"ShareLab-SII","fullname":"ShareLab-SII","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d4b07fa99b113364b9ca86/nTBRWVJ08tHVh-mh6i8tL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.18607.md","query":{}}">
Papers
arxiv:2608.18607

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Published on Aug 19
· Submitted by
Yinming Huang
on Aug 20
Authors:
,

Abstract

A human-aligned chain-of-thought reward model and preference dataset improve joint video-audio generation by replacing fragmented metrics with coherent, dimension-wise reinforcement learning.

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.

Community

🚀 We are excited to introduce VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation, the first reward model specifically designed for joint video-audio generation.

Existing metrics typically evaluate video and audio separately, overlooking the holistic cross-modal coherence that shapes human preference. When used for post-training, these fragmented metrics may also lead to reward hacking, where metric scores improve without a corresponding improvement in perceptual quality.

Our key contributions include:

  • VA-Judger, a reasoning-based omni-modal reward model that jointly assesses visual and audio quality, text alignment, audio-video synchronization, semantic coherence, and overall human preference.
  • VAPref-10K and VA-Judger-Bench, which provide human preference annotations and a challenging benchmark covering both in-domain and out-of-domain video-audio generation models.
  • A complete reward-modeling and post-training framework that uses VA-Judger to improve joint video-audio generation.

VA-Judger substantially outperforms single-dimensional metrics and omni-modal model baselines such as Qwen3-Omni. It also generalizes reliably to unseen closed-source generation models.

When used to post-train LTX-2, the resulting model achieves a 62.30% human preference rate, compared with 27.63% for the OmniNFT-trained version and 10.08% for the original LTX-2. It also achieves the best performance on 11 out of 13 objective metrics.

🔗 Project: https://sharelab-sii.github.io/VA-Judger/
📄 Paper: https://arxiv.org/abs/2608.18607
💻 Code: https://github.com/ShareLab-SII/VA-Judger
🤗 Models: https://huggingface.co/ShareLab-SII/VA-Judger
📊 Dataset: https://huggingface.co/datasets/ShareLab-SII/VA-Judger-Bench
🎮 Demo: https://www.youtube.com/watch?v=HUiEFLTY9-E

Further training code and the full VAPref-10K dataset will be released soon. Stay tuned!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.18607
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.18607 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.18607 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.18607 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers