Hugging Face Daily Papers · · 4 min read

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

FIRM-Video introduces a checklist-driven, check-before-score framework for reliable and efficient text-to-video reward modeling. It decomposes evaluation into verifiable criteria for instruction following, world coherence, and perceptual quality, and builds FIRM-Video-90K and FIRM-Video-Bench. FIRM-Video-8B achieves the best overall MAE on the benchmark and consistently improves Best-of-8 video selection across multiple generators.</p>\n","updatedAt":"2026-08-27T05:14:35.370Z","author":{"_id":"6710be3e6d1b33cf24417e38","avatarUrl":"/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg","fullname":"zpy","name":"zpy777","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8034195303916931},"editors":["zpy777"],"editorAvatarUrls":["/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.21839","authors":[{"_id":"6a8fc6663bd48bb654ea68ce","name":"Peiyuan Zhang","hidden":false},{"_id":"6a8fc6663bd48bb654ea68cf","name":"Xiangyu Zhao","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d0","name":"Hongbo Liu","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d1","name":"Xiaoxing Hu","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d2","name":"Mingxin Liu","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d3","name":"Shuran Ma","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d4","name":"Yunhang Shen","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d5","name":"Jian Hu","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d6","name":"Haihan Gao","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d7","name":"Haoyu Cao","hidden":false},{"_id":"6a8fc6663bd48bb654ea68d8","name":"Xue Yang","hidden":false}],"publishedAt":"2026-08-22T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling","submittedOnDailyBy":{"_id":"6710be3e6d1b33cf24417e38","avatarUrl":"/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg","isPro":false,"fullname":"zpy","user":"zpy777","type":"user","name":"zpy777"},"summary":"Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.","upvotes":1,"discussionId":"6a8fc6673bd48bb654ea68d9","projectPage":"https://firm-reward.github.io/","ai_summary":"FIRM-Video uses checklist-driven verification of temporal visual evidence to build reliable reward models for text-to-video evaluation and alignment.","ai_keywords":["reward models","text-to-video","checklist-driven","check-before-score","instruction following","world coherence","perceptual quality","temporal visual evidence","reward modeling","FIRM-Video-8B"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6867f748a177ff0ff64431e0","name":"VisionXLab","fullname":"SJTU VisionXLab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648e77184cae4f6921dbb382/8Fmw6rFukqZo5SYhuHcPw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6710be3e6d1b33cf24417e38","avatarUrl":"/avatars/f60bc9a67bb58f5997cbcc28cb93c079.svg","isPro":false,"fullname":"zpy","user":"zpy777","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6867f748a177ff0ff64431e0","name":"VisionXLab","fullname":"SJTU VisionXLab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648e77184cae4f6921dbb382/8Fmw6rFukqZo5SYhuHcPw.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.21839.md","query":{}}">
Papers
arxiv:2608.21839

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

Published on Aug 22
· Submitted by
zpy
on Aug 27
Authors:
,

Abstract

FIRM-Video uses checklist-driven verification of temporal visual evidence to build reliable reward models for text-to-video evaluation and alignment.

Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Community

Paper submitter about 4 hours ago

FIRM-Video introduces a checklist-driven, check-before-score framework for reliable and efficient text-to-video reward modeling. It decomposes evaluation into verifiable criteria for instruction following, world coherence, and perceptual quality, and builds FIRM-Video-90K and FIRM-Video-Bench. FIRM-Video-8B achieves the best overall MAE on the benchmark and consistently improves Best-of-8 video selection across multiple generators.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.21839
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.21839 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.21839 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.21839 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers