Hugging Face Daily Papers · · 3 min read

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

An easy-to-use toolkit for fine-grained robot process assessment, with progress-aware metrics, progress judge benchmarking, and visualization tools.</p>\n","updatedAt":"2026-08-17T04:09:38.993Z","author":{"_id":"668f5478b3991ac0c3fc9c2f","avatarUrl":"/avatars/a775853d3b88e7b1c8494ca837b5495c.svg","fullname":"yuhengji","name":"yuheng2000","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8599610924720764},"editors":["yuheng2000"],"editorAvatarUrls":["/avatars/a775853d3b88e7b1c8494ca837b5495c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.14284","authors":[{"_id":"6a82830fb601d59c65281454","name":"Yuyang Liu","hidden":false},{"_id":"6a82830fb601d59c65281455","name":"Yanqing Shen","hidden":false},{"_id":"6a82830fb601d59c65281456","name":"Ruike Chen","hidden":false},{"_id":"6a82830fb601d59c65281457","name":"Jifan Zhao","hidden":false},{"_id":"6a82830fb601d59c65281458","name":"Yuxuan Tian","hidden":false},{"_id":"6a82830fb601d59c65281459","name":"Yichi Zhang","hidden":false},{"_id":"6a82830fb601d59c6528145a","name":"Tianfeng Long","hidden":false},{"_id":"6a82830fb601d59c6528145b","name":"Zixuan Yin","hidden":false},{"_id":"6a82830fb601d59c6528145c","name":"Yipu Wang","hidden":false},{"_id":"6a82830fb601d59c6528145d","name":"Ziheng Qin","hidden":false},{"_id":"6a82830fb601d59c6528145e","name":"Wenxing Tan","hidden":false},{"_id":"6a82830fb601d59c6528145f","name":"Yang Shi","hidden":false},{"_id":"6a82830fb601d59c65281460","name":"Mingyu Cao","hidden":false},{"_id":"6a82830fb601d59c65281461","name":"Runze Xiao","hidden":false},{"_id":"6a82830fb601d59c65281462","name":"Ziqi Wang","hidden":false},{"_id":"6a82830fb601d59c65281463","name":"Zhixin Yin","hidden":false},{"_id":"6a82830fb601d59c65281464","name":"Shiwei Chu","hidden":false},{"_id":"6a82830fb601d59c65281465","name":"Yi-Fan Zhang","hidden":false},{"_id":"6a82830fb601d59c65281466","name":"Yao Mu","hidden":false},{"_id":"6a82830fb601d59c65281467","name":"Yuheng Ji","hidden":false},{"_id":"6a82830fb601d59c65281468","name":"Yihao Wang","hidden":false},{"_id":"6a82830fb601d59c65281469","name":"Jun Yan","hidden":false},{"_id":"6a82830fb601d59c6528146a","name":"Zhongyuan Wang","hidden":false},{"_id":"6a82830fb601d59c6528146b","name":"Pengwei Wang","hidden":false},{"_id":"6a82830fb601d59c6528146c","name":"Xiaolong Zheng","hidden":false}],"publishedAt":"2026-08-14T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment","submittedOnDailyBy":{"_id":"668f5478b3991ac0c3fc9c2f","avatarUrl":"/avatars/a775853d3b88e7b1c8494ca837b5495c.svg","isPro":false,"fullname":"yuhengji","user":"yuheng2000","type":"user","name":"yuheng2000"},"summary":"Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.","upvotes":6,"discussionId":"6a82830fb601d59c6528146d","projectPage":"https://prm-as-a-judge.github.io/","ai_summary":"PRM-as-a-Judge 1.5 provides fine-grained process metrics and reliability tools to evaluate embodied robotic models beyond binary success rates.","ai_keywords":["process reward models","rollout videos","progress curves","failure-side progress","post-drawdown recovery","execution quality","RoboPulse++","embodied models"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"668f5478b3991ac0c3fc9c2f","avatarUrl":"/avatars/a775853d3b88e7b1c8494ca837b5495c.svg","isPro":false,"fullname":"yuhengji","user":"yuheng2000","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"67ea560810394e02bce763ed","avatarUrl":"/avatars/870df00bc66b10685a1bae1c70e0b87c.svg","isPro":false,"fullname":"lyy","user":"lyy0715","type":"user"},{"_id":"671fa0de0a34a05602094909","avatarUrl":"/avatars/e35890e1f71a8c3f4e33b35a027e180a.svg","isPro":false,"fullname":"Yanqing Shen","user":"syq1105","type":"user"},{"_id":"6901e42c8611d14fa3a0bd19","avatarUrl":"/avatars/00e17f840ea40b75a4b443e6e7060813.svg","isPro":false,"fullname":"ltf","user":"ltf101","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.14284.md","query":{}}">
Papers
arxiv:2608.14284

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Published on Aug 14
· Submitted by
yuhengji
on Aug 17
Authors:
,

Abstract

PRM-as-a-Judge 1.5 provides fine-grained process metrics and reliability tools to evaluate embodied robotic models beyond binary success rates.

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

Community

Paper submitter about 4 hours ago

An easy-to-use toolkit for fine-grained robot process assessment, with progress-aware metrics, progress judge benchmarking, and visualization tools.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.14284
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.14284 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.14284 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.14284 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers