ChatPaper.aiChatPaper

GigaWorld-1: 로봇 정책 평가를 위한 세계 모델 구축 로드맵

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

July 2, 2026
저자: GigaWorld Team, Angyuan Ma, Boyuan Wang, Bohan Li, Chaojun Ni, Guo Li, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jingyu Liu, Jiwen Lu, Qiuping Deng, Tingdong Yu, Xuancheng Xu, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Xiaofeng Wang, Xiaoyu Tian, Yang Wang, Yifan Chang, Yukun Zhou, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu
cs.AI

초록

구현된 로봇 기반 모델을 평가하는 것은 여전히 중요한 병목 현상으로 남아 있다. 디지털 벤치마크를 통해 효율적으로 평가되는 대규모 언어 모델과 달리, 로봇 정책은 하드웨어와 인간의 감독에 의해 제한되는 느리고 비용이 많이 드는 실제 롤아웃을 필요로 하며, 이는 대리 정책 평가자로서 세계 모델에 대한 관심을 불러일으켰다. 그러나 정책 평가를 위해 세계 모델을 신뢰할 수 있게 만드는 핵심 속성은 여전히 제대로 이해되지 않고 있다. 본 연구는 로봇 정책 평가를 위한 세계 모델에 대한 체계적인 연구를 제시하고, WMBench를 소개한다. WMBench는 실제 로봇 원격 조작 데이터와 일치하는 정책 롤아웃으로 구성된 벤치마크로, 다양한 조작 작업을 포괄하여 모델 군, 행동 인코딩, 롤아웃 수평선 및 평가 지표 간 통제된 비교를 가능하게 한다. WMBench를 사용하여 7개의 비디오 세계 모델, 4가지 행동 표현 방식, 실제 로봇 실행과 짝을 이루는 324,000개 이상의 시뮬레이션 정책 롤아웃을 분석하며, CVPR 2026 GigaBrain Challenge의 대규모 커뮤니티 제출물, 선별된 합성 궤적, 12,000시간이 넘는 훈련 비디오를 추가하여 분석을 더욱 풍부하게 한다. 실험을 통해 세 가지 핵심 통찰을 얻었다. 평가자 품질은 단기적 시각적 사실성보다 장기적이고 행동에 충실한 롤아웃 일관성에 의해 좌우된다. 사전 훈련의 이점은 데이터 규모뿐만 아니라 일반적인 세계 지식과 로봇 특화 제어 가능성 간의 균형에서 비롯된다. 행동 인코딩, 메모리 설계, 평가자 중심 사후 훈련을 포함한 아키텍처 선택은 실제 로봇 행동과의 정렬을 강하게 결정한다. 이러한 결과를 바탕으로 실용적인 설계 로드맵을 도출하고, 정책 평가에 특별히 최적화된 세계 모델인 GigaWorld-1에서 이를 구현한다. 구현된 기반 모델을 위한 확장 가능한 평가 연구를 발전시키기 위해 코드, 모델, 데이터 세트 및 도구 키트를 완전히 공개한다.
English
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in GigaWorld-1, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.