GigaWorld-Policy-0.5: AutoResearchによって強化された、より高速で強力なWAM
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
July 15, 2026
著者: GigaWorld Team, Angen Ye, Angyuan Ma, Boyuan Wang, Chaojun Ni, Fangzheng Ye, Guan Huang, Guo Li, Guosheng Zhao, Haodong Yan, Hengtao Li, Jiwen Lu, Kai Wang, Mingming Yu, Qitang Hu, Qiuping Deng, Songling Liu, Xiaoyu Tian, Xiaofeng Wang, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Yang Wang, Yejun Zeng, Yifan Chang, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu
cs.AI
要旨
ワールドアクションモデル(WAM)は、動作と将来の視覚観測を同時にモデル化し、将来のシーン変化を物理的に根拠づけられた動作生成のための密な教師信号として利用することで、ロボットの政策学習を改善する。しかし、既存のWAMにおける一般的な設計では、推論時に将来の動画を明示的に生成する必要があり、これにより大幅な計算オーバーヘッドが生じ、リアルタイムの閉ループ展開が妨げられる。GigaWorld-Policyは、訓練時には将来の視覚ダイナミクスを活用しつつ、推論時には動作のみのデコードを行う動作中心の定式化により、この問題に対処する。本フレームワークに基づき、より効率的なロボット制御を目的とした強化版動作中心WAMであるGigaWorld-Policy-0.5を提案する。事前学習では、GigaWorld-Policy-0.5は混合されたAction-Conditioned World Modeling(AC-WM)とWAM訓練戦略を採用する。これにより、視覚ダイナミクスとロボット動作の間の結合が強化され、下流の政策学習における動作表現の転移可能性が向上する。効率的な推論のために、GigaWorld-Policy-0.5はMixture-of-Transformersアーキテクチャを導入し、視覚ダイナミクスのモデル化と動作生成を専門化されたエキスパートに分離する。これにより、動作のみの推論時のアクティブ計算が削減され、ローカルRTX 4090環境において85msの推論遅延を達成する。さらに、エージェントベースのAutoResearchパイプラインを活用して訓練設定を系統的に探索し、ハイパーパラメータ調整に要する時間と手動介入を低減しつつ、最適な実験設定をより効率的に特定する。実験およびアブレーション研究により、GigaWorld-Policy-0.5は将来の視覚ダイナミクスの訓練上の利点を保持しつつ、ロボット制御における推論効率を向上させることが示される。
English
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual dynamics are used during training while action-only decoding is used at inference time. Building upon this framework, we present GigaWorld-Policy-0.5, an enhanced action-centered WAM designed for more efficient robot control. During pretraining, GigaWorld-Policy-0.5 adopts a mixed Action-Conditioned World Modeling (AC-WM) and WAM training strategy. This strengthens the coupling between visual dynamics and robot actions and improves the transferability of action representations for downstream policy learning. For efficient inference, GigaWorld-Policy-0.5 introduces a Mixture-of-Transformers architecture that separates visual dynamics modeling and action generation into specialized experts, reducing active computation during action-only inference and achieving 85 ms inference latency on a local RTX 4090 setup. In addition, we employ an agent-based AutoResearch pipeline to systematically search training configurations, enabling more efficient identification of optimal experimental setups while reducing the time and manual intervention required for hyperparameter tuning. Experiments and ablations show that GigaWorld-Policy-0.5 preserves the training benefits of future visual dynamics while improving inference efficiency for robot control.