ChatPaper.aiChatPaper

PRM-as-a-Judge 1.5: 로봇 프로세스 평가를 위한 도구 키트

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

August 14, 2026
저자: Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao, Runze Xiao, Ziqi Wang, Zhixin Yin, Shiwei Chu, Yi-Fan Zhang, Yao Mu, Yuheng Ji, Yihao Wang, Jun Yan, Zhongyuan Wang, Pengwei Wang, Xiaolong Zheng
cs.AI

초록

세밀한 로봇 평가는 이진 성공률과 규칙 기반 프로세스 점수를 넘어 체화된 모델을 이해하는 데 중요하다. 우리는 롤아웃 비디오를 조밀한 진행 곡선으로 변환하고 여러 세밀한 지표를 도출하는 로봇 프로세스 평가 도구 키트인 PRM-as-a-Judge 1.5를 제시한다. PRM-as-a-Judge 1.5는 버전 1.0을 기반으로, 실패 측 진행, 하락 후 회복, 성공 측 실행 품질을 특성화하는 세 가지 지표를 도입하여 사용자가 체화된 모델의 능력을 이해하도록 돕는다. 벤치마크의 롤아웃 비디오를 기반으로 우리는 체화된 모델에 대한 포괄적인 평가를 수행하고, 세밀한 지표 결과와 주요 발견을 제공한다. 또한 프로세스 보상 모델(PRM)의 신뢰성을 평가하기 위해 RoboPulse++를 도입하여 평가자에게 더 정확한 테스트 플랫폼을 제공한다. 더욱이, 우리는 벤치마크, 지표 구현 및 시각화 도구를 포함한 사용자 친화적인 평가 스위트를 공개하여 재현 가능한 조작 프로세스 평가를 지원한다. 우리는 커뮤니티가 로봇 평가 방식을 재고하고, 투명하고 절차적이며 재현 가능한 평가를 차세대 체화된 지능의 기반으로 확립할 것을 촉구한다.
English
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.