Hugging Face Daily Papers · · 4 min read

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Homepage: <a href=\"https://video-reason.com/\" rel=\"nofollow\">https://video-reason.com/</a><br>Data and Models: <a href=\"https://huggingface.co/collections/Video-Reason/vbvr-pro-a-scalable-and-verifiable-suite-for-native-visual\">https://huggingface.co/collections/Video-Reason/vbvr-pro-a-scalable-and-verifiable-suite-for-native-visual</a><br>Training Code: <a href=\"https://github.com/Video-Reason/VBVR-Pro\" rel=\"nofollow\">https://github.com/Video-Reason/VBVR-Pro</a><br>Eval Code: <a href=\"https://github.com/Video-Reason/VBVR-Pro-Bench\" rel=\"nofollow\">https://github.com/Video-Reason/VBVR-Pro-Bench</a></p>\n","updatedAt":"2026-08-27T03:12:04.689Z","author":{"_id":"652d06833b5997ed71ce5c46","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/O_D6bpa5mGxLA7uCjmVCG.jpeg","fullname":"Zhongang Cai","name":"caizhongang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":39,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/k66xcOMf4NVbMSFulUjHY.png","fullname":"SenseNova","name":"sensenova","type":"org","isHf":false,"plan":"team"}}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.6389989256858826},"editors":["caizhongang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/O_D6bpa5mGxLA7uCjmVCG.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26105","authors":[{"_id":"6a8f9afb2c24e8c5fab328d9","name":"Junxiang Xu","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328da","name":"Ruisi Wang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328db","user":{"_id":"646e1ef5075bbcc48ddf21e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646e1ef5075bbcc48ddf21e8/g-nFu-plmEdTpnAJh_pUx.png","isPro":false,"fullname":"Pu Fanyi","user":"pufanyi","type":"user","name":"pufanyi"},"name":"Fanyi Pu","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.920Z","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328dc","name":"Maijunxian Wang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328dd","name":"Ran Ji","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328de","name":"Tongxi Zhou","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328df","name":"Chenyang Gu","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e0","name":"Jing Zuo","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e1","name":"Hongcan Xiao","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e2","name":"Yimeng Geng","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e3","name":"Wanqi Yin","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e4","name":"Wei Chen","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e5","name":"Oscar Qian","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e6","name":"Zhengan Yan","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e7","name":"Ziqi Huang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e8","user":{"_id":"64b4a717aa03b6520839e9b8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b4a717aa03b6520839e9b8/Rt3ERG-6BVEA4hAwOz0_I.jpeg","isPro":false,"fullname":"Haiwen Diao","user":"Paranioar","type":"user","name":"Paranioar"},"name":"Haiwen Diao","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.914Z","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328e9","name":"Liang Pan","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328ea","name":"Bo Li","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328eb","name":"Xiangyu Fan","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328ec","name":"Dezhi Luo","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328ed","name":"Fengyuan Yu","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328ee","name":"Zehong Zhao","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328ef","name":"Qingying Gao","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f0","name":"Tinghui Zhu","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f1","name":"Yilan Zhang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f2","name":"Jingqi Tong","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f3","name":"Pinyuan Feng","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f4","name":"Zhengze Jiang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f5","name":"Letian Wang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f6","name":"Ziyu Guo","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f7","name":"Renrui Zhang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f8","name":"Jieneng Chen","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328f9","name":"Sonia Joseph","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328fa","name":"Constantin Venhoff","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328fb","name":"Saman Motamed","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328fc","name":"Mengyue Yang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328fd","name":"Chandra Sripada","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328fe","name":"Alan Yuille","hidden":false},{"_id":"6a8f9afb2c24e8c5fab328ff","name":"Philip Torr","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32900","name":"Lvmin Zhang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32901","name":"Vikash Kumar","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32902","name":"Daniel Khashabi","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32903","name":"Nikolaus Kriegeskorte","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32904","name":"Raphaël Millière","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32905","name":"Vincent C. Müller","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32906","name":"Anyi Rao","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32907","name":"Quan Wang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32908","name":"Ziwei Liu","hidden":false},{"_id":"6a8f9afb2c24e8c5fab32909","name":"Dahua Lin","hidden":false},{"_id":"6a8f9afb2c24e8c5fab3290a","name":"Lei Yang","hidden":false},{"_id":"6a8f9afb2c24e8c5fab3290b","name":"Hokin Deng","hidden":false},{"_id":"6a8f9afb2c24e8c5fab3290c","user":{"_id":"652d06833b5997ed71ce5c46","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/O_D6bpa5mGxLA7uCjmVCG.jpeg","isPro":false,"fullname":"Zhongang Cai","user":"caizhongang","type":"user","name":"caizhongang"},"name":"Zhongang Cai","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.908Z","hidden":false}],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.","upvotes":17,"discussionId":"6a8f9afb2c24e8c5fab3290d","ai_summary":"VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates.","ai_keywords":["native visual reasoning","visual generation","MLLMs","VLM-as-a-judge","verifiable reward scorers","multi-task reinforcement learning","interleaved generation","vision-native trajectories"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"652d06833b5997ed71ce5c46","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652d06833b5997ed71ce5c46/O_D6bpa5mGxLA7uCjmVCG.jpeg","isPro":false,"fullname":"Zhongang Cai","user":"caizhongang","type":"user"},{"_id":"62ab1ac1d48b4d8b048a3473","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1656826685333-62ab1ac1d48b4d8b048a3473.png","isPro":false,"fullname":"Ziwei Liu","user":"liuziwei7","type":"user"},{"_id":"67f87529318a17cc80365190","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f87529318a17cc80365190/kv4cAvD5BrWQFRXKG4FXg.jpeg","isPro":false,"fullname":"Maijunxian Wang","user":"Mark7121983123","type":"user"},{"_id":"68139a6caa391fe39ff3143f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/-FpdbEoaGFaBgasBsTzy1.png","isPro":false,"fullname":"Quan Wang","user":"wangquan12","type":"user"},{"_id":"668f51fb8e8d87dbdd23caa9","avatarUrl":"/avatars/0b17cb0f3ad2c729f185cdccdad94e48.svg","isPro":false,"fullname":"Yin Wanqi","user":"waanqii","type":"user"},{"_id":"6562db047061c2cbda120f02","avatarUrl":"/avatars/ac1a917a99e1291e751a68091ae678be.svg","isPro":false,"fullname":"Ketone","user":"Olefine","type":"user"},{"_id":"644e3e5f030210812f413073","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uW8TKV2sds97lFBwnt6JK.jpeg","isPro":true,"fullname":"Zilong Chen","user":"heheyas","type":"user"},{"_id":"64b4a717aa03b6520839e9b8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b4a717aa03b6520839e9b8/Rt3ERG-6BVEA4hAwOz0_I.jpeg","isPro":false,"fullname":"Haiwen Diao","user":"Paranioar","type":"user"},{"_id":"669f13ed48d82d4a6cfc367b","avatarUrl":"/avatars/84fe7eacdaef02e4e161bfdb5b59cd71.svg","isPro":false,"fullname":"Jing Zuo","user":"Panda7777777","type":"user"},{"_id":"66c7360df375ce3a32dd9fa0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/elGBNDjw3WeV8ZEjgQ_HA.png","isPro":false,"fullname":"Zhe Cao","user":"MichaelCaoo","type":"user"},{"_id":"646e1ef5075bbcc48ddf21e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646e1ef5075bbcc48ddf21e8/g-nFu-plmEdTpnAJh_pUx.png","isPro":false,"fullname":"Pu Fanyi","user":"pufanyi","type":"user"},{"_id":"680b0d0f1173808fedf31f96","avatarUrl":"/avatars/ec920acf1200fc301737529aaabc87db.svg","isPro":false,"fullname":"Junxiang Xu","user":"junxiangjoe","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26105.md","query":{}}">
Papers
arxiv:2608.26105

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Published on Aug 26
· Submitted by
taesiri
on Aug 27
Authors:
,

Abstract

VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates.

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.26105
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Browse 15 models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.26105 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers