ChatPaper.aiChatPaper

SG-WAM: 幾何学を考慮したポリシー空間における自己誘導型世界モデリング

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

August 2, 2026
著者: Ruiteng Zhao, Zhengshen Zhang, Yue Su, Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang Jr., Haiyue Zhu
cs.AI

要旨

世界行動モデル(WAMs)は、行動生成と将来状態の予測を結合する。その有効性は、将来のダイナミクスが、行動生成と整合的であり、かつ行動がシーンをどこでどのように変化させるかを捉えるのに十分な幾何学的認識を備えた空間においてモデル化されるかどうかに依存する。既存のWAMは通常、この要件の一部しか満たしておらず、知覚的に重い観測空間のターゲットか、行動関連性と幾何学のために共同で構造化されていない補助的な潜在空間のいずれかに依存している。本稿では、方策由来の表現空間において幾何学的認識を備えた行動条件付きダイナミクスを直接学習する自己導出型フレームワークSG-WAMを提案する。SG-WAMは、学習可能なダイナミクストークンと、介在するロボットの行動を条件としてそれらの将来の潜在状態を予測する自己導出型世界予測器(Self-Guided World Predictor)を導入する。予測ターゲットは、同一の方策バックボーンの指数移動平均コピーによって生成され、行動エキスパートが使用する表現ファミリー内で安定した教師信号を提供する。幾何学的教師信号は、方策の画像トークン表現をさらに構造化し、ダイナミクストークンに空間的に基盤づけられたコンテキストを提供して、行動関連性があり幾何学的認識を備えた将来整合空間を生み出す。潜在空間における将来予測、幾何学的基盤付け、およびフローマッチング行動生成は、統一フレームワーク内でエンドツーエンドに共同最適化される。大規模な身体性事前学習を伴わない0.9Bモデルに基づき、SG-WAMはLIBEROで平均98.5%の成功率、LIBERO-Plusで73%を達成し、分布内および分布外の実世界評価の両方において強力なベースラインを上回る。
English
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and how actions change the scene. Existing WAMs typically satisfy only part of this requirement, relying on either perceptually heavy observation-space targets or auxiliary latent spaces that are not jointly structured for action relevance and geometry. We propose SG-WAM, a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space. SG-WAM introduces learnable dynamics tokens and a Self-Guided World Predictor that forecasts their future latent states conditioned on intervening robot actions. Prediction targets are generated by an exponential moving average copy of the same policy backbone, providing stable supervision within the representation family used by the action expert. Geometric supervision further structures the policy image-token representations, providing spatially grounded context for the dynamics tokens and yielding a future-alignment space that is both action-relevant and geometry-aware. Latent future prediction, geometric grounding, and flow-matching action generation are jointly optimized end-to-end in a unified framework. Built on a 0.9B model without large-scale embodied pretraining, SG-WAM achieves 98.5% average success on LIBERO and 73% on LIBERO-Plus, while outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.