ChatPaper.aiChatPaper

RoboDojo: 범용 로봇 조작 정책의 종합 평가를 위한 통합 시뮬-실제 벤치마크

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

July 7, 2026
저자: Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Haoran Lu, Weijie Wan, Baijun Chen, Songling Liu, Haowen Yan, Honghao Su, Zhiyang Dou, Kaixuan Wang, Dandan Zhang, Yunze Liu, Yan Qin, Qiwei Liang, Qiwei Wu, Zijian Lin, Wenwei Lin, Yuran Wang, Minghua He, Tianshu Wu, Ruihai Wu, Jingquan Zhou, Kai-Chong Lei, Haibao Yu, Yuanfeng Ji, Weiyang Jin, Guanyu Lin, Xiaofan Li, Qi Xiong, Renjing Xu, Zhongyu Li, Wenhao Chai, Enze Xie, Ziwei Wang, Yao Mu, Hao Dong, Wojciech Matusik, Mingyu Ding, Wenbo Ding, Ping Luo, Masayoshi Tomizuka
cs.AI

초록

범용 로봇 조작 정책(Generalist Robot Manipulation Policies)은 빠르게 발전하고 있지만, 기존 벤치마크는 이들의 성능을 체계적으로 평가하는 데 있어 여전히 한계가 있습니다. 많은 벤치마크가 단순하고 단기적이거나 기술 범위가 좁은 과제에 의존하며, 보통 시뮬레이션 또는 실제 환경 중 한쪽에서만 수행됩니다. 시뮬레이션은 확장 가능한 피드백을 제공하지만 물리적 배포 과정의 어려움을 반영하지 못하는 반면, 실제 환경 평가는 비용이 많이 들고 시간이 오래 소요되며 재현이 어렵습니다. 본 연구에서는 범용 로봇 조작 정책의 포괄적인 평가를 위한 통합형 시뮬레이션-실물 벤치마크인 RoboDojo를 소개합니다. RoboDojo는 42개의 시뮬레이션 과제와 18개의 실제 환경 과제로 구성되어 있으며, 이들은 다양하고 상호 보완적인 조작 능력을 포괄합니다. 시뮬레이션 벤치마크는 일반화, 기억력, 정밀도, 장기 과제 수행, 개방형 어휘 명령 수행의 다섯 가지 차원을 평가하고, 실제 환경 벤치마크는 정책을 까다로운 물리적 배포 조건에 노출시킵니다. RoboDojo는 Isaac Sim 내 이기종 병렬 시뮬레이션을 통한 확장 가능한 평가를 지원하며, 원격 클라우드 접근, 표준화된 하드웨어, 장면 재설정, 평가 프로토콜 및 배포 인터페이스를 갖춘 재현 가능한 실제 환경 평가 시스템인 RoboDojo-RealEval을 제공합니다. XPolicyLab과 함께, 정책은 한 번 통합되어 최소한의 수정만으로 시뮬레이션 및 실제 환경에서 평가할 수 있습니다. 우리는 30개의 정책을 XPolicyLab에 통합하고 RoboDojo에서 평가하여 공개 리더보드와 현재 정책 성능에 대한 체계적인 분석을 구축했습니다. 웹사이트는 http://robodojo-benchmark.com/에서 확인할 수 있습니다.
English
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.