SG-WAM: 기하학 인식 정책 공간에서의 자기 유도 세계 모델링
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
August 2, 2026
저자: Ruiteng Zhao, Zhengshen Zhang, Yue Su, Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang Jr., Haiyue Zhu
cs.AI
초록
세계 행동 모델(World Action Models, WAM)은 행동 생성과 미래 상태 예측을 결합한다. 이들의 효과성은 미래 역학이 행동 생성과 정렬되는 동시에, 행동이 장면에서 일으키는 변화의 위치와 방식을 포착할 수 있을 만큼 기하학 인식적인 공간에서 모델링되는지에 달려 있다. 기존 WAM은 일반적으로 이러한 요구사항의 일부만을 충족하며, 지각적 부담이 큰 관찰 공간 목표나 행동 관련성과 기하학적 구조를 위해 공동으로 구조화되지 않은 보조 잠재 공간에 의존한다. 본 논문은 정책 파생 표현 공간에서 기하학 인식 행동 조건 역학을 직접 학습하는 자기 유도 프레임워크인 SG-WAM을 제안한다. SG-WAM은 학습 가능한 역학 토큰과, 개입하는 로봇 행동에 조건화된 미래 잠재 상태를 예측하는 자기 유도 세계 예측기(Self-Guided World Predictor)를 도입한다. 예측 목표는 동일한 정책 백본의 지수 이동 평균(EMA) 사본으로 생성되어, 행동 전문가가 사용하는 표현 계열 내에서 안정적인 감독을 제공한다. 기하학적 감독은 정책 이미지 토큰 표현을 추가로 구조화하여 역학 토큰에 공간적으로 근거된 맥락을 부여하고, 행동 관련성과 기하학 인식성을 모두 갖춘 미래 정렬 공간을 산출한다. 잠재 미래 예측, 기하학적 그라운딩, 플로우 매칭 행동 생성은 통합 프레임워크에서 종단 간 공동 최적화된다. 대규모 체화 사전학습 없이 0.9B 모델을 기반으로 구축된 SG-WAM은 LIBERO에서 평균 98.5%, LIBERO-Plus에서 73%의 성공률을 달성하며, 분포 내(in-distribution) 및 분포 외(out-of-distribution) 실세계 평가 모두에서 강력한 베이스라인들을 능가한다.
English
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and how actions change the scene. Existing WAMs typically satisfy only part of this requirement, relying on either perceptually heavy observation-space targets or auxiliary latent spaces that are not jointly structured for action relevance and geometry. We propose SG-WAM, a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space. SG-WAM introduces learnable dynamics tokens and a Self-Guided World Predictor that forecasts their future latent states conditioned on intervening robot actions. Prediction targets are generated by an exponential moving average copy of the same policy backbone, providing stable supervision within the representation family used by the action expert. Geometric supervision further structures the policy image-token representations, providing spatially grounded context for the dynamics tokens and yielding a future-alignment space that is both action-relevant and geometry-aware. Latent future prediction, geometric grounding, and flow-matching action generation are jointly optimized end-to-end in a unified framework. Built on a 0.9B model without large-scale embodied pretraining, SG-WAM achieves 98.5% average success on LIBERO and 73% on LIBERO-Plus, while outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.