We introduce VGI-Bench, a benchmark for probing visual intelligence in video generation models beyond perceptual quality. It evaluates whether video generators can actually reason through evolving visual processes, with 27 tasks and 810 instances covering diverse reasoning skills. Our results show that current models exhibit promising visual reasoning capabilities, but remain far from reliable—even the strongest evaluated model reaches only 51.0%. We hope VGI-Bench can help reveal where video generation models truly reason, where they fail, and what is needed for the next generation of visually intelligent models.</p>\n","updatedAt":"2026-08-27T02:47:47.693Z","author":{"_id":"655c3953ce055ed40a7de0ba","avatarUrl":"/avatars/70aad195f6aff83279ef1b1f8859419d.svg","fullname":"Xuan He","name":"hexuan21","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.901260495185852},"editors":["hexuan21"],"editorAvatarUrls":["/avatars/70aad195f6aff83279ef1b1f8859419d.svg"],"reactions":[{"reaction":"🚀","users":["CongWei1230","AgPerry"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.19583","authors":[{"_id":"6a8a90623d26296ea3091620","name":"Xuan He","hidden":false},{"_id":"6a8a90623d26296ea3091621","user":{"_id":"64f8e358766ff9f3d2b0de84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f8e358766ff9f3d2b0de84/R2P1YG-mRBh7TU9wkjGGk.jpeg","isPro":true,"fullname":"Cong Wei","user":"CongWei1230","type":"user","name":"CongWei1230"},"name":"Cong Wei","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.832Z","hidden":false},{"_id":"6a8a90623d26296ea3091622","name":"Yuhao Cheng","hidden":false},{"_id":"6a8a90623d26296ea3091623","name":"Linrui Ma","hidden":false},{"_id":"6a8a90623d26296ea3091624","user":{"_id":"63b908d0e3c78740d8e950d0","avatarUrl":"/avatars/3e80075e92aebdfea712f70b00d5ec7d.svg","isPro":false,"fullname":"Yuxuan Zhang","user":"Reacherx","type":"user","name":"Reacherx"},"name":"Yuxuan Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.839Z","hidden":false},{"_id":"6a8a90623d26296ea3091625","name":"Zuojun Li","hidden":false},{"_id":"6a8a90623d26296ea3091626","name":"Yuhao Wen","hidden":false},{"_id":"6a8a90623d26296ea3091627","name":"Zeyi Liu","hidden":false},{"_id":"6a8a90623d26296ea3091628","name":"Yuren Hao","hidden":false},{"_id":"6a8a90623d26296ea3091629","name":"Songcheng Cai","hidden":false},{"_id":"6a8a90623d26296ea309162a","name":"Keming Wu","hidden":false},{"_id":"6a8a90623d26296ea309162b","name":"Penghui Du","hidden":false},{"_id":"6a8a90623d26296ea309162c","name":"Kai Zou","hidden":false},{"_id":"6a8a90623d26296ea309162d","name":"Rui Yang","hidden":false},{"_id":"6a8a90623d26296ea309162e","name":"Chenkai Sun","hidden":false},{"_id":"6a8a90623d26296ea309162f","name":"Ke Yang","hidden":false},{"_id":"6a8a90623d26296ea3091630","name":"Ping Nie","hidden":false},{"_id":"6a8a90623d26296ea3091631","name":"Kelsey R Allen","hidden":false},{"_id":"6a8a90623d26296ea3091632","name":"Chenglong Wang","hidden":false},{"_id":"6a8a90623d26296ea3091633","name":"Michel Galley","hidden":false},{"_id":"6a8a90623d26296ea3091634","name":"Jianfeng Gao","hidden":false},{"_id":"6a8a90623d26296ea3091635","name":"ChengXiang Zhai","hidden":false}],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"VGI-BENCH: Probing Visual Intelligence in Video Generation Models","submittedOnDailyBy":{"_id":"655c3953ce055ed40a7de0ba","avatarUrl":"/avatars/70aad195f6aff83279ef1b1f8859419d.svg","isPro":true,"fullname":"Xuan He","user":"hexuan21","type":"user","name":"hexuan21"},"summary":"Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance~2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. We will release our code and data.","upvotes":61,"discussionId":"6a8a90633d26296ea3091636","projectPage":"https://hexuan21.github.io/VGI-Bench/","githubRepo":"https://github.com/hexuan21/VGI-Bench","githubRepoAddedBy":"user","ai_summary":"VGI-bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and minimal self-correction during generation.","ai_keywords":["video generation models","zero-shot visual reasoning","VGI-bench","visual reasoning capabilities","failure modes","input condition sensitivity","synthetic fine-tuning","denoising","self-correction"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"655c3953ce055ed40a7de0ba","avatarUrl":"/avatars/70aad195f6aff83279ef1b1f8859419d.svg","isPro":true,"fullname":"Xuan He","user":"hexuan21","type":"user"},{"_id":"65358802a920f38780b3248a","avatarUrl":"/avatars/9415510b598079973c2b0436ad12db9c.svg","isPro":false,"fullname":"Ping Nie","user":"pingnieuk","type":"user"},{"_id":"6981745aa12f5a5f980b2e6c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/s7_SLjrFPCUDJszvHh4h0.webp","isPro":false,"fullname":"Zuojun Li","user":"pear-tree","type":"user"},{"_id":"670d7aed372cb8fadbd270bb","avatarUrl":"/avatars/652c8b90d492e3e5db6c735d69ae0991.svg","isPro":false,"fullname":"Songcheng Cai","user":"SongchengCai","type":"user"},{"_id":"64f8e358766ff9f3d2b0de84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f8e358766ff9f3d2b0de84/R2P1YG-mRBh7TU9wkjGGk.jpeg","isPro":true,"fullname":"Cong Wei","user":"CongWei1230","type":"user"},{"_id":"62290ea25860cd00e2600fd0","avatarUrl":"/avatars/3dd440775cd06e318115fc7fb5157e6f.svg","isPro":false,"fullname":"Huang","user":"Chengyu","type":"user"},{"_id":"662a7759efa616e734ab493d","avatarUrl":"/avatars/8b79c6ec01f13d6b82414a6ad1b2d588.svg","isPro":false,"fullname":"Ke Yang","user":"EmpathYang","type":"user"},{"_id":"691e5f168f82cd99d66df74d","avatarUrl":"/avatars/ca614ef49da66cdb3f1d5ac07118ed9f.svg","isPro":true,"fullname":"Perry the Platypus","user":"AgPerry","type":"user"},{"_id":"63b908d0e3c78740d8e950d0","avatarUrl":"/avatars/3e80075e92aebdfea712f70b00d5ec7d.svg","isPro":false,"fullname":"Yuxuan Zhang","user":"Reacherx","type":"user"},{"_id":"6923a72078e21722ab0363ef","avatarUrl":"/avatars/9c9410cb31201648dd5fcacb6ac0080e.svg","isPro":false,"fullname":"Linrui Ma","user":"JerryLin828","type":"user"},{"_id":"6a8519b16295d50f571724d2","avatarUrl":"/avatars/edfcad0a92ff3c778081bac9a45bcee0.svg","isPro":false,"fullname":"Bruno Costa","user":"brunocost","type":"user"},{"_id":"6a80de808d1455f75e991be1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a80de808d1455f75e991be1/LQzyEe-eiqY3nQ34OwK6H.jpeg","isPro":false,"fullname":"Dmitri Pavlov","user":"dpavlov0623","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.19583.md","query":{}}">
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Abstract
VGI-bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and minimal self-correction during generation.
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance~2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. We will release our code and data.
Community
We introduce VGI-Bench, a benchmark for probing visual intelligence in video generation models beyond perceptual quality. It evaluates whether video generators can actually reason through evolving visual processes, with 27 tasks and 810 instances covering diverse reasoning skills. Our results show that current models exhibit promising visual reasoning capabilities, but remain far from reliable—even the strongest evaluated model reaches only 51.0%. We hope VGI-Bench can help reveal where video generation models truly reason, where they fail, and what is needed for the next generation of visually intelligent models.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.19583 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.19583 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.