OpenWAM:邁向系統性世界-動作模型預訓練的開放式模組化探索
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
September 7, 2026
作者: Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao
cs.AI
摘要
世界-動作模型從影片生成先驗繼承世界知識,並透過具身經驗將其轉化為可執行的控制訊號。然而,現有系統是單體式的:生成主幹、視覺表徵、架構、資訊流、推論程序與訓練資料緊密耦合,模糊了哪些設計選擇重要及其原因。我們提出 OpenWAM,一個開放研究堆疊,將世界-動作預訓練轉化為受控實驗計畫。OpenWAM-Infra 將 WAM 設計空間分解為可組合模組,並具備統一的訓練、推論、部署與評估。在此基底上,OpenWAM-Study 透過受控實驗探討三個問題:該繼承什麼、世界與動作學習如何互動,以及其協同效應如何擴展;並提煉三項原則:上游知識透過足夠強大的生成主幹與緊湊且資訊豐富的潛在空間進行轉移;世界-動作協同需要專用動作能力、明確的世界到動作資訊流,以及同步聯合去噪;而具身預訓練主要提升域外泛化,並透過對第一人稱視角與機器人資料進行單階段共同訓練,整合世界覆蓋與動作接地。綜合這些原則,我們建構了 OpenWAM-α,一個開放的 WAM,在約 6,400 小時的第一人稱視角人類與機器人資料上預訓練,並在模擬與真實世界基準上評估。在八個模擬基準與真實機器人實驗中,這些實驗共同涵蓋從單臂與雙臂操作到靈巧手的具身形式,OpenWAM-α 展現一致優異的效能,從模擬到實體世界維持其頂尖地位。我們釋出完整堆疊,包括基礎設施、評估協議、預訓練模型與資料配方,以促進未來研究。
English
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.