Hugging Face Daily Papers · · 5 min read

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

HarnessEval is an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval interprets the context of each evaluation case, decomposes the evaluation question into measurable sub-questions, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own sub-question. The parent agent then validates the gathered evidence and aggregates it into the final verdict. Every evaluation becomes a transparent evidence tree whose complete reasoning chain justifies the result.</p>\n","updatedAt":"2026-08-18T03:54:19.193Z","author":{"_id":"643b866bff50448bcfc7d1d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/q12v-BatVKimoi8q-coi-.jpeg","fullname":"Jialong Wu","name":"manchery","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8814935088157654},"editors":["manchery"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/q12v-BatVKimoi8q-coi-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16859","authors":[{"_id":"6a83cc8e675db694db8cd4ed","user":{"_id":"6578007d9fae206bdfa5ee2b","avatarUrl":"/avatars/7bc86a68562bdabfa497d57abd8642f2.svg","isPro":false,"fullname":"Weiliang Chen","user":"chen-wl20","type":"user","name":"chen-wl20"},"name":"Weiliang Chen","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.163Z","hidden":false},{"_id":"6a83cc8e675db694db8cd4ee","name":"Haowen Sun","hidden":false},{"_id":"6a83cc8e675db694db8cd4ef","user":{"_id":"633b7a4b0d68f86e2d98de05","avatarUrl":"/avatars/5d48c171ddbcc7ca39bdc0d11c6224e4.svg","isPro":false,"fullname":"Jun Gao","user":"JungaoCanada","type":"user","name":"JungaoCanada"},"name":"Jun Gao","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.156Z","hidden":false},{"_id":"6a83cc8e675db694db8cd4f0","name":"Jiawei Chi","hidden":false},{"_id":"6a83cc8e675db694db8cd4f1","name":"Hanyang Wang","hidden":false},{"_id":"6a83cc8e675db694db8cd4f2","name":"Qiyu Dai","hidden":false},{"_id":"6a83cc8e675db694db8cd4f3","name":"Yihao Li","hidden":false},{"_id":"6a83cc8e675db694db8cd4f4","name":"Hao Li","hidden":false},{"_id":"6a83cc8e675db694db8cd4f5","name":"Jingnan Gao","hidden":false},{"_id":"6a83cc8e675db694db8cd4f6","name":"Yi-Hsin Hung","hidden":false},{"_id":"6a83cc8e675db694db8cd4f7","name":"Xingzhuo Guo","hidden":false},{"_id":"6a83cc8e675db694db8cd4f8","name":"Shangchen Miao","hidden":false},{"_id":"6a83cc8e675db694db8cd4f9","name":"Zhiyuan Shi","hidden":false},{"_id":"6a83cc8e675db694db8cd4fa","name":"Xiang Li","hidden":false},{"_id":"6a83cc8e675db694db8cd4fb","name":"Fengrui Tian","hidden":false},{"_id":"6a83cc8e675db694db8cd4fc","user":{"_id":"66f080bc789ce1b57b605379","avatarUrl":"/avatars/0f164d69802b257881a5f23ca10a4726.svg","isPro":false,"fullname":"Weihua Du","user":"VanishD","type":"user","name":"VanishD"},"name":"Weihua Du","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.177Z","hidden":false},{"_id":"6a83cc8e675db694db8cd4fd","name":"Ziqi Huang","hidden":false},{"_id":"6a83cc8e675db694db8cd4fe","name":"Shenyuan Gao","hidden":false},{"_id":"6a83cc8e675db694db8cd4ff","name":"Siqiao Huang","hidden":false},{"_id":"6a83cc8e675db694db8cd500","name":"Mingyu Liu","hidden":false},{"_id":"6a83cc8e675db694db8cd501","name":"Yifei Li","hidden":false},{"_id":"6a83cc8e675db694db8cd502","name":"Shizun Wang","hidden":false},{"_id":"6a83cc8e675db694db8cd503","name":"Xi Wang","hidden":false},{"_id":"6a83cc8e675db694db8cd504","name":"Tianqi Zhang","hidden":false},{"_id":"6a83cc8e675db694db8cd505","name":"Xue Luo","hidden":false},{"_id":"6a83cc8e675db694db8cd506","name":"Xiyin Ren","hidden":false},{"_id":"6a83cc8e675db694db8cd507","name":"Jinshan Ren","hidden":false},{"_id":"6a83cc8e675db694db8cd508","name":"Xiaoyang Shen","hidden":false},{"_id":"6a83cc8e675db694db8cd509","name":"Xiaobo Hu","hidden":false},{"_id":"6a83cc8e675db694db8cd50a","name":"Zhiyang Dou","hidden":false},{"_id":"6a83cc8e675db694db8cd50b","name":"Mingyu Ding","hidden":false},{"_id":"6a83cc8e675db694db8cd50c","name":"Yichao Yan","hidden":false},{"_id":"6a83cc8e675db694db8cd50d","name":"Xinchao Wang","hidden":false},{"_id":"6a83cc8e675db694db8cd50e","name":"Yizhou Wang","hidden":false},{"_id":"6a83cc8e675db694db8cd50f","name":"Shilong Liu","hidden":false},{"_id":"6a83cc8e675db694db8cd510","name":"Wenzhao Zheng","hidden":false},{"_id":"6a83cc8e675db694db8cd511","name":"Yueqi Duan","hidden":false},{"_id":"6a83cc8e675db694db8cd512","name":"Yuan Gong","hidden":false},{"_id":"6a83cc8e675db694db8cd513","name":"Ziwei Liu","hidden":false},{"_id":"6a83cc8e675db694db8cd514","name":"Ming-Yu Liu","hidden":false},{"_id":"6a83cc8e675db694db8cd515","user":{"_id":"643b866bff50448bcfc7d1d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/q12v-BatVKimoi8q-coi-.jpeg","isPro":true,"fullname":"Jialong Wu","user":"manchery","type":"user","name":"manchery"},"name":"Jialong Wu","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.169Z","hidden":false},{"_id":"6a83cc8e675db694db8cd516","name":"Jiangran Lyu","hidden":false},{"_id":"6a83cc8e675db694db8cd517","name":"Fangfu Liu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/643b866bff50448bcfc7d1d1/NYWxBcKynYoSRrHS--99s.mp4"],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"HarnessEval-W: Agentifying the Evaluation of Visual Worlds","submittedOnDailyBy":{"_id":"643b866bff50448bcfc7d1d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/q12v-BatVKimoi8q-coi-.jpeg","isPro":true,"fullname":"Jialong Wu","user":"manchery","type":"user","name":"manchery"},"summary":"A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.","upvotes":50,"discussionId":"6a83cc8e675db694db8cd518","projectPage":"https://mirros-lab.github.io/HarnessEval-W","githubRepo":"https://github.com/MirroS-Lab/HarnessEval-W","githubRepoAddedBy":"user","ai_summary":"HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence.","ai_keywords":["HarnessEval-W","agentified evaluation pipeline","harness paradigm","world model benchmarking","sub-agents","evidence tree","reasoning chain"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":14,"organization":{"_id":"6a77dec42fad4e89f1ce346a","name":"MirroS-Lab","fullname":"MirroS","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/643b866bff50448bcfc7d1d1/pSoXy9fqBTNK3AUuisaH-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6578007d9fae206bdfa5ee2b","avatarUrl":"/avatars/7bc86a68562bdabfa497d57abd8642f2.svg","isPro":false,"fullname":"Weiliang Chen","user":"chen-wl20","type":"user"},{"_id":"643b866bff50448bcfc7d1d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/q12v-BatVKimoi8q-coi-.jpeg","isPro":true,"fullname":"Jialong Wu","user":"manchery","type":"user"},{"_id":"6577fd0c2c8d6e12c4df5865","avatarUrl":"/avatars/83d264d01bc21765952225ddd1c1a389.svg","isPro":false,"fullname":"Haowen Sun","user":"Sunboy0617","type":"user"},{"_id":"675163177679a2657e6677e8","avatarUrl":"/avatars/e2c4aeeb615ce84d4e9e0b78bd205736.svg","isPro":false,"fullname":"Xingzhuo Guo","user":"GXZlegend","type":"user"},{"_id":"6463245c4ad7e61e51db0b2a","avatarUrl":"/avatars/ea7937c94a84c7ec255cd47ed5161aa9.svg","isPro":false,"fullname":"Jingnan Gao","user":"G1nOnly","type":"user"},{"_id":"69b53c14c88f9126b4ccf317","avatarUrl":"/avatars/8534e30890bf31672c14e7ba3b685883.svg","isPro":false,"fullname":"Ruichen Wang","user":"wrc02","type":"user"},{"_id":"641bc0057c21ab946bf63c1b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641bc0057c21ab946bf63c1b/aHUHlR4X-0k0WFjpxoyYi.png","isPro":false,"fullname":"Zhiyuan Shi","user":"2304022380szy","type":"user"},{"_id":"6784c13588796724ed982d61","avatarUrl":"/avatars/76bf00a17af9cb8aebe0dac000de5f9b.svg","isPro":false,"fullname":"Yingyue Li","user":"liyy4586","type":"user"},{"_id":"67ea02f6048f3bf1dd56f3b4","avatarUrl":"/avatars/0abfc0ec090b28ebfd6b6a4d7abb74d0.svg","isPro":false,"fullname":"Jiangran Lyu","user":"Jiangranlyu","type":"user"},{"_id":"66e14cf2fb009ef598305fe5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/etohfcD9GKaRPir1HRP-5.jpeg","isPro":false,"fullname":"Jiawei Chi","user":"chijw","type":"user"},{"_id":"673ad2e6a852d37889d53c94","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/EeU0BonSi-qng7Qw-veri.png","isPro":false,"fullname":"Yuezhou Ma","user":"Pharosmyz0609","type":"user"},{"_id":"6407501f5e6d06cc2cf5dd13","avatarUrl":"/avatars/6df2d0fb01901dffdd35d0b7723e5867.svg","isPro":false,"fullname":"沈俊涛","user":"suly233333","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"6a77dec42fad4e89f1ce346a","name":"MirroS-Lab","fullname":"MirroS","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/643b866bff50448bcfc7d1d1/pSoXy9fqBTNK3AUuisaH-.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16859.md","query":{}}">
Papers
arxiv:2608.16859

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Published on Aug 17
· Submitted by
Jialong Wu
on Aug 18
#1 Paper of the day

Abstract

HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence.

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

Community

Paper author Paper submitter about 5 hours ago

HarnessEval is an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval interprets the context of each evaluation case, decomposes the evaluation question into measurable sub-questions, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own sub-question. The parent agent then validates the gathered evidence and aggregates it into the final verdict. Every evaluation becomes a transparent evidence tree whose complete reasoning chain justifies the result.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16859
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16859 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16859 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16859 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers