OpenWAM: 체계적인 월드-액션 모델 사전학습을 향한 개방형 모듈식 탐구
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
September 7, 2026
저자: Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao
cs.AI
초록
월드-액션 모델은 비디오 생성 사전으로부터 월드 지식을 계승하고, 이를 체화된 경험을 통해 실행 가능한 제어 신호로 전환한다. 그러나 기존 시스템은 모놀리식하여, 생성 백본, 시각적 표현, 아키텍처, 정보 흐름, 추론 절차, 학습 데이터가 긴밀하게 결합되어 있어 어떤 설계 선택이 중요하고 그 이유가 무엇인지 불분명하게 만든다. 우리는 월드-액션 사전학습을 통제된 실험 프로그램으로 전환하는 개방형 연구 스택인 OpenWAM을 제안한다. OpenWAM-Infra는 WAM 설계 공간을 통합된 학습, 추론, 배포, 평가를 갖춘 조합 가능한 모듈로 분해한다. 이 기반 위에서 OpenWAM-Study는 통제된 실험을 통해 무엇을 계승할지, 월드 학습과 행동 학습이 어떻게 상호작용하는지, 그리고 이들의 시너지가 어떻게 확장되는지라는 세 가지 질문을 검토하고, 세 가지 원리를 도출한다: 업스트림 지식은 충분히 강력한 생성 백본과 컴팩트하고 정보가 풍부한 잠재 공간을 통해 전이된다; 월드-액션 시너지는 전용 행동 용량, 월드에서 행동으로의 명시적 정보 흐름, 동기화된 공동 디노이징을 요구한다; 그리고 체화 사전학습은 주로 도메인 외 일반화를 향상시키며, 에고센트릭 인간 및 로봇 데이터에 대한 단일 단계 공동 학습은 월드 커버리지와 행동 그라운딩을 통합한다. 이러한 원리를 구성하여, 우리는 약 6,400시간 분량의 에고센트릭 인간 및 로봇 데이터로 사전학습되고 시뮬레이션 및 실제 세계 벤치마크 전반에서 평가된 개방형 WAM인 OpenWAM-α를 구축한다. 단일 팔 및 양팔 조작에서 정교한 로봇 손에 이르기까지 구현체를 함께 아우르는 8개의 시뮬레이션 벤치마크와 실제 로봇 실험 전반에서, OpenWAM-α는 일관되게 뛰어난 성능을 제공하며 시뮬레이션에서 물리적 세계에 이르기까지 최상위 위치를 유지한다. 우리는 향후 연구를 촉진하기 위해 인프라, 평가 프로토콜, 사전학습 모델, 데이터 레시피를 포함한 전체 스택을 공개한다.
English
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.