ZimaBlue: 확장 가능한 비디오 사전 학습을 통한 일반화 가능한 세계 행동 모델의 진화
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
August 31, 2026
저자: Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun, Jianhui Liu, Songsong Yu, Jiyao Zhang, Wenbo Li, Bo Wang, Guoqing Ma, Lin Song, Renjie Liao, Shenghe Zheng, Wei Tang, Xiaojuan Qi, Yanwei Li, Yuan Zhang, Zhuotao Tian, Haoyang Huang, Nan Duan
cs.AI
초록
로봇 조작은 근본적인 확장(scaling) 문제에 직면해 있다. 강건한 일반화는 광범위한 물리적 경험을 요구하지만, 행동 레이블이 포함된 로봇 궤적은 수집 비용이 높고 본질적으로 다양성이 제한적이다. 1인칭(egocentric) 영상은 다양한 환경에서의 객체 상호작용, 접촉 역학, 도구 사용, 장기 행동을 포착함으로써 훨씬 더 확장 가능한 체화된(embodied) 경험의 원천을 제공한다. 핵심 과제는 이처럼 풍부하지만 행동 정보가 없는 경험을 효과적인 로봇 제어로 변환하는 것이다. 본 논문에서는 대규모 영상으로부터 일반화 가능한 세계 행동 모델(World Action Model, WAM)을 학습하는 확장 가능한 프레임워크인 ZimaBlue를 제안한다. ZimaBlue는 3단계 학습 커리큘럼을 따른다. 먼저 대규모 인간 및 로봇 1인칭 영상에 대한 인과적 체화 영상 사전학습(causal embodied video pre-training)을 수행한다. 다음으로, 통합된 행동 표현을 이용한 비디오-행동 중간 학습(video-action mid-training)을 통해 학습된 시각적 역학을 이종 로봇 궤적에 접지(grounding)한다. 마지막으로 실제 배포를 위해 모델을 대상 로봇에 특화시킨다. 생성형 WAM을 실시간 제어에 실용적으로 사용하기 위해 ZimaBlue는 비동기식 Slow-Fast 이중 시스템 아키텍처를 추가로 채택한다. 고용량의 Slow 세계 모델은 일반화 가능한 시공간 표현을 제공하고, 경량 Fast 분기는 NVIDIA RTX 4090에서 30Hz의 행동 예측을 가능하게 한다. 실물 로봇 제로샷(zero-shot) 평가에서 학습 데이터를 대상 로봇 궤적만으로 한정한 설정에서 120,000시간이 넘는 체화 영상까지 확장했을 때 성공률은 36.1%에서 77.8%로 향상되었다. ZimaBlue는 또한 여러 벤치마크에서 강력한 성능을 보여주며, 특히 본 적 없는 작업(unseen tasks)에서 성능 향상이 두드러진다.
English
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of embodied experience, capturing object interactions, contact dynamics, tool use, and long-horizon behaviors across diverse environments. The central challenge is how to convert this abundant but action-free experience into effective robot control. We introduce ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video. ZimaBlue follows a three-stage training curriculum: it first performs causal embodied video pre-training on large-scale human and robot egocentric videos, then grounds the learned visual dynamics in heterogeneous robot trajectories through video-action mid-training with a unified action representation, and finally specializes the model to a target robot for deployment. To make generative WAMs practical for real-time control, ZimaBluefurther adopts an asynchronous Slow-Fast dual-system architecture, where a high-capacity Slow world model provides generalizable spatiotemporal representations and a lightweight Fast branch enables 30 Hz action prediction on NVIDIA RTX 4090. On real-robot zero-shot evaluations, scaling from target-robot data alone to over 120,000 hours of embodied video improves success from 36.1% to 77.8%. ZimaBlue further delivers strong performance across multiple benchmarks, with particularly pronounced gains on unseen tasks.