ChatPaper.aiChatPaper

GigaWorld-1: ロボットポリシー評価のための世界モデル構築ロードマップ

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

July 2, 2026
著者: GigaWorld Team, Angyuan Ma, Boyuan Wang, Bohan Li, Chaojun Ni, Guo Li, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jingyu Liu, Jiwen Lu, Qiuping Deng, Tingdong Yu, Xuancheng Xu, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Xiaofeng Wang, Xiaoyu Tian, Yang Wang, Yifan Chang, Yukun Zhou, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu
cs.AI

要旨

具現化ロボット基盤モデルの評価は依然として重要なボトルネックである。大規模言語モデルがデジタルベンチマークにより効率的に評価できるのとは異なり、ロボットポリシーはハードウェアや人間の監視に制約された低速でコストのかかる実機展開を必要とする。このため、代理ポリシー評価器としてのワールドモデルへの関心が高まっているが、ポリシー評価においてワールドモデルを信頼できるものにするための主要な特性は未だ十分に理解されていない。本研究では、ロボットポリシー評価のためのワールドモデルに関する体系的な研究を提示し、実機ロボットの遠隔操作データとそれに対応するポリシーロールアウトから構成されるベンチマークWMBenchを導入する。WMBenchは多様な操作タスクをカバーし、モデルファミリ、行動エンコーディング、ロールアウト期間、評価指標にわたる制御比較を可能にする。WMBenchを用いて、7つのビデオワールドモデル、4つの行動表現手法、および実機ロボットの実行と対になる32万4千以上のシミュレーションポリシーロールアウトを分析する。さらに、CVPR 2026 GigaBrain Challengeからの大規模コミュニティ投稿、厳選された合成軌道、および12,000時間を超える訓練動画を用いて分析を充実させる。実験から得られた3つの核心的な知見は以下のとおりである。評価器の品質は、短期的な視覚的リアリズムよりも、長期的で行動に忠実なロールアウトの一貫性によって決まる。事前学習の利得はデータ規模だけでなく、一般的な世界知識とロボット固有の制御可能性のバランスから生じる。行動エンコーディング、メモリ設計、評価器重視の追加学習などのアーキテクチャ上の選択が、実機ロボットの動作との整合性を強く決定する。これらの結果に基づき、実用的な設計ロードマップを導出し、それをポリシー評価に特化して最適化されたワールドモデルGigaWorld-1として具現化する。また、コード、モデル、データセット、ツールキットを完全に公開し、具現化基盤モデルのスケーラブルな評価研究を推進する。
English
Evaluating embodied robot foundation models remains a critical bottleneck; unlike large language models efficiently assessed via digital benchmarks, robotic policies require slow, costly real-world rollouts limited by hardware and human supervision, which has driven interest in world models as surrogate policy evaluators, yet the key properties that make a world model reliable for policy assessment remain poorly understood. This work presents a systematic study of world models for robotic policy evaluation and introduces WMBench, a benchmark constructed from real-robot teleoperation data and matched policy rollouts covering diverse manipulation tasks to enable controlled comparisons across model families, action encodings, rollout horizons, and evaluation metrics. Using WMBench, we analyze 7 video world models, 4 action representation schemes, and over 324,000 simulated policy rollouts paired with real robot executions, further enriching our analysis with large-scale community submissions from the CVPR 2026 GigaBrain Challenge, curated synthetic trajectories, and a training videos spanning more than 12,000 hours. Our experiments deliver three core insights: evaluator quality is dominated by long-horizon, action-faithful rollout consistency rather than short-term visual realism; pretraining gains stem not only from data scale but from balancing general world knowledge with robot-specific controllability; and architectural choices including action encoding, memory design, and evaluator-focused post-training strongly determine alignment with real-world robot behavior. Drawing on these results, we derive a practical design roadmap and realize it in GigaWorld-1, a world model specially optimized for policy evaluation, and we fully release our code, models, datasets, and toolkits to advance scalable evaluation research for embodied foundation models.