PRM-as-a-Judge 1.5:ロボットプロセス評価のためのツールキット
PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
August 14, 2026
著者: Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao, Runze Xiao, Ziqi Wang, Zhixin Yin, Shiwei Chu, Yi-Fan Zhang, Yao Mu, Yuheng Ji, Yihao Wang, Jun Yan, Zhongyuan Wang, Pengwei Wang, Xiaolong Zheng
cs.AI
要旨
きめ細かいロボット評価は、二値的な成功率やルールベースのプロセススコアを超えて、具現化モデルを理解するために重要です。我々は、ロールアウト動画を密な進捗曲線に変換し、複数の詳細な指標を導出するロボットプロセス評価のためのツールキットであるPRM-as-a-Judge 1.5を提示します。PRM-as-a-Judge 1.5は、バージョン1.0を基に、失敗側の進捗、ドローダウン後の回復、成功側の実行品質を特徴付ける3つの指標を導入し、ユーザーが具現化モデルの能力を理解するのに役立ちます。ベンチマークからのロールアウト動画に基づき、我々は具現化モデルの包括的な評価を実施し、いくつかの詳細な指標結果と重要な発見を提供します。さらに、プロセス報酬モデル(PRM)の信頼性を評価するためにRoboPulse++を導入し、評価者により正確なテストプラットフォームを提供します。加えて、ベンチマーク、指標の実装、可視化ツールを含むユーザーフレンドリーな評価スイートを公開し、再現可能な操作プロセス評価を支援します。我々はコミュニティに対し、ロボットの評価方法を再考し、透明で手続き的かつ再現可能な評価を次世代の具現化知能の基盤として確立するよう呼びかけます。
English
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.