GigaWorld-Policy-0.5: AutoResearch로 강화된 더욱 빠르고 강력한 WAM
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
July 15, 2026
저자: GigaWorld Team, Angen Ye, Angyuan Ma, Boyuan Wang, Chaojun Ni, Fangzheng Ye, Guan Huang, Guo Li, Guosheng Zhao, Haodong Yan, Hengtao Li, Jiwen Lu, Kai Wang, Mingming Yu, Qitang Hu, Qiuping Deng, Songling Liu, Xiaoyu Tian, Xiaofeng Wang, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Yang Wang, Yejun Zeng, Yifan Chang, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu
cs.AI
초록
세계 행동 모델(World Action Models, WAMs)은 행동과 미래 시각 관측을 공동으로 모델링하고, 미래 장면의 변화를 물리적 기반 행동 생성을 위한 조밀한 감독 신호로 사용하여 로봇 정책 학습을 향상시킨다. 그러나 기존 WAM의 일반적인 설계는 추론 시 미래 비디오를 명시적으로 생성하므로 상당한 계산 오버헤드가 발생하고 실시간 폐쇄 루프 배포를 방해한다. GigaWorld-Policy는 행동 중심 공식화를 통해 이 문제를 해결하며, 훈련 중에는 미래 시각 역학을 사용하고 추론 시에는 행동 전용 디코딩을 사용한다. 이 프레임워크를 기반으로, 우리는 보다 효율적인 로봇 제어를 위해 설계된 향상된 행동 중심 WAM인 GigaWorld-Policy-0.5를 제시한다. 사전 학습 중 GigaWorld-Policy-0.5는 혼합된 행동 조건부 세계 모델링(Action-Conditioned World Modeling, AC-WM) 및 WAM 학습 전략을 채택한다. 이를 통해 시각 역학과 로봇 행동 간의 결합을 강화하고 하위 정책 학습을 위한 행동 표현의 전이 가능성을 향상시킨다. 효율적인 추론을 위해 GigaWorld-Policy-0.5는 Mixture-of-Transformers 아키텍처를 도입하여 시각 역학 모델링과 행동 생성을 전문화된 전문가로 분리함으로써, 행동 전용 추론 중 활성 계산을 줄이고 로컬 RTX 4090 설정에서 85ms의 추론 지연 시간을 달성한다. 또한, 에이전트 기반 AutoResearch 파이프라인을 사용하여 훈련 구성을 체계적으로 탐색함으로써, 하이퍼파라미터 튜닝에 필요한 시간과 수동 개입을 줄이면서 최적의 실험 설정을 보다 효율적으로 식별할 수 있다. 실험 및 절제 연구는 GigaWorld-Policy-0.5가 로봇 제어를 위한 추론 효율성을 향상시키면서 미래 시각 역학의 훈련 이점을 유지함을 보여준다.
English
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual dynamics are used during training while action-only decoding is used at inference time. Building upon this framework, we present GigaWorld-Policy-0.5, an enhanced action-centered WAM designed for more efficient robot control. During pretraining, GigaWorld-Policy-0.5 adopts a mixed Action-Conditioned World Modeling (AC-WM) and WAM training strategy. This strengthens the coupling between visual dynamics and robot actions and improves the transferability of action representations for downstream policy learning. For efficient inference, GigaWorld-Policy-0.5 introduces a Mixture-of-Transformers architecture that separates visual dynamics modeling and action generation into specialized experts, reducing active computation during action-only inference and achieving 85 ms inference latency on a local RTX 4090 setup. In addition, we employ an agent-based AutoResearch pipeline to systematically search training configurations, enabling more efficient identification of optimal experimental setups while reducing the time and manual intervention required for hyperparameter tuning. Experiments and ablations show that GigaWorld-Policy-0.5 preserves the training benefits of future visual dynamics while improving inference efficiency for robot control.