ChatPaper.aiChatPaper

OpenWAM:体系的なWorld-Action Model事前学習に向けたオープンでモジュール型の探索

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

September 7, 2026
著者: Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao
cs.AI

要旨

World-Actionモデルは、動画生成事前分布から世界知識を継承し、身体化経験を通じてそれを実行可能な制御信号へと導く。しかしながら、既存システムはモノリシックであり、生成バックボーン、視覚表現、アーキテクチャ、情報フロー、推論手順、訓練データが密結合しているため、どの設計選択が重要で、なぜ重要かが不明瞭になっている。我々は、World-Action事前学習を制御された実験プログラムへと変えるオープン研究スタックであるOpenWAMを導入する。OpenWAM-Infraは、WAMの設計空間を、統合された訓練・推論・デプロイ・評価を備えた構成可能なモジュールへと因子分解する。この基盤上で、OpenWAM-Studyは制御実験を通じて、何を継承するか、世界学習と行動学習がどのように相互作用するか、それらの相乗効果がどのようにスケールするかという3つの問いを検討し、次の3つの原理を抽出する。すなわち、上流知識は、十分に高性能な生成バックボーンと、コンパクトで情報豊富な潜在空間を通じて転移する;World-Actionの相乗効果には、専用の行動容量、明示的な世界から行動への情報フロー、同期された共同ノイズ除去が必要である;そして、身体化事前学習は主にドメイン外汎化を改善し、エゴセントリックデータとロボットデータに対する一段階の共同訓練が、世界のカバレッジと行動のグラウンディングを統合する。これらの原理を組み合わせて、我々はOpenWAM-αを構築する。これは、約6,400時間のエゴセントリックな人間・ロボットデータで事前学習され、シミュレーションおよび実世界ベンチマークで評価されるオープンなWAMである。8つのシミュレーションベンチマークと実ロボット実験にわたり、これらは合わさって単腕および両腕マニピュレーションから多指ハンドに至るまでの身体性を網羅するが、OpenWAM-αは一貫して優れた性能を発揮し、シミュレーションから物理世界に至るまで最高水準の地位を維持する。今後の研究を促進するため、インフラストラクチャ、評価プロトコル、事前学習済みモデル、データレシピを含むフルスタックを公開する。
English
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.