GigaWorld-Policy-0.5:由AutoResearch賦能的速度更快、功能更強的WAM
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
July 15, 2026
作者: GigaWorld Team, Angen Ye, Angyuan Ma, Boyuan Wang, Chaojun Ni, Fangzheng Ye, Guan Huang, Guo Li, Guosheng Zhao, Haodong Yan, Hengtao Li, Jiwen Lu, Kai Wang, Mingming Yu, Qitang Hu, Qiuping Deng, Songling Liu, Xiaoyu Tian, Xiaofeng Wang, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Yang Wang, Yejun Zeng, Yifan Chang, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu
cs.AI
摘要
世界行動模型(WAMs)透過聯合建模動作與未來視覺觀測,利用未來場景演化作為密集監督信號,以提升機器人策略學習的物理基礎動作生成能力。然而,現有WAMs的常見設計是在推理階段明確生成未來影片,導致龐大的計算負擔,阻礙即時閉環部署的可行性。GigaWorld-Policy 以動作中心架構解決此問題:訓練階段利用未來視覺動態,但推理階段僅進行動作解碼。在此框架基礎上,我們提出 GigaWorld-Policy-0.5——一種強化的動作中心 WAM,專為更高效的機器人控制而設計。預訓練階段,GigaWorld-Policy-0.5 採用混合式動作條件世界模型(AC-WM)與 WAM 訓練策略,強化視覺動態與機器人動作之間的耦合性,並提升動作表徵於下游策略學習的可遷移性。為實現高效推理,GigaWorld-Policy-0.5 引入混合專家Transformer架構,將視覺動態建模與動作生成分離為專業化專家,減少僅進行動作推論時的活躍計算量,並在本地 RTX 4090 環境下達成 85 毫秒推理延遲。此外,我們採用基於智能代理的自動化研究(AutoResearch)管線,系統性搜尋訓練配置,在縮短超參數調優所需時間與人工介入的同時,更有效地識別最佳實驗設置。實驗與消融研究結果顯示,GigaWorld-Policy-0.5 在保留未來視覺動態訓練效益的同時,提升了機器人控制的推論效率。
English
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual dynamics are used during training while action-only decoding is used at inference time. Building upon this framework, we present GigaWorld-Policy-0.5, an enhanced action-centered WAM designed for more efficient robot control. During pretraining, GigaWorld-Policy-0.5 adopts a mixed Action-Conditioned World Modeling (AC-WM) and WAM training strategy. This strengthens the coupling between visual dynamics and robot actions and improves the transferability of action representations for downstream policy learning. For efficient inference, GigaWorld-Policy-0.5 introduces a Mixture-of-Transformers architecture that separates visual dynamics modeling and action generation into specialized experts, reducing active computation during action-only inference and achieving 85 ms inference latency on a local RTX 4090 setup. In addition, we employ an agent-based AutoResearch pipeline to systematically search training configurations, enabling more efficient identification of optimal experimental setups while reducing the time and manual intervention required for hyperparameter tuning. Experiments and ablations show that GigaWorld-Policy-0.5 preserves the training benefits of future visual dynamics while improving inference efficiency for robot control.