RoboDojo:統一虛實整合基準,用於全面評估通用型機器人操作策略
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
July 7, 2026
作者: Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Haoran Lu, Weijie Wan, Baijun Chen, Songling Liu, Haowen Yan, Honghao Su, Zhiyang Dou, Kaixuan Wang, Dandan Zhang, Yunze Liu, Yan Qin, Qiwei Liang, Qiwei Wu, Zijian Lin, Wenwei Lin, Yuran Wang, Minghua He, Tianshu Wu, Ruihai Wu, Jingquan Zhou, Kai-Chong Lei, Haibao Yu, Yuanfeng Ji, Weiyang Jin, Guanyu Lin, Xiaofan Li, Qi Xiong, Renjing Xu, Zhongyu Li, Wenhao Chai, Enze Xie, Ziwei Wang, Yao Mu, Hao Dong, Wojciech Matusik, Mingyu Ding, Wenbo Ding, Ping Luo, Masayoshi Tomizuka
cs.AI
摘要
通用型机器人操控策略近年来发展迅速,但现有基准测试在系统评估其能力方面仍存在局限。许多基准测试依赖于简单、短时程或技能单一的任务,能力覆盖范围有限,且通常仅在仿真或现实环境中进行。仿真虽能实现可扩展的反馈,但忽略了物理部署的实际挑战;而现实环境评估成本高、耗时长且难以复现。为此,我们提出RoboDojo——一个统一的模拟与真实基准测试,用于全面评估通用型机器人操控策略。RoboDojo包含42个仿真任务与18个现实任务,覆盖多样化且互补的操控能力。仿真基准从五个维度进行评估:泛化性、记忆能力、精度、长时程执行以及开放式指令跟随;而现实基准则将策略置于具有挑战性的物理世界部署条件下。RoboDojo通过Isaac Sim中的异构并行仿真支持可扩展的评估,并提供RoboDojo-RealEval——一个可复现的现实评估系统,具备远程云端访问、标准化硬件、场景重置、评估协议及部署接口。结合XPolicyLab,策略可一次性集成,并在仅需最小调整的情况下在模拟与现实环境中进行评估。我们将30种策略集成至XPolicyLab,并在RoboDojo上对其进行评估,建立了公开排行榜与当前策略性能的系统性分析。相关网站详见http://robodojo-benchmark.com/。
English
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.