ChatPaper.aiChatPaper

ST-WAM: 視覚的分布シフト下でのロバストなマニピュレーションのための意味的・時間的ワールドアクションモデル

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

July 31, 2026
著者: Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li
cs.AI

要旨

World Action Models (WAMs)は、ロボットの行動と将来の視覚的ダイナミクスを統合的にモデル化する有望なパラダイムとして登場している。しかし、そのピクセル生成型の将来予測への依存は、行動に関連する状態遷移とタスクに無関係な視覚コンテンツを絡み合わせる可能性があり、視覚的分布シフト下での堅牢性を制限している。我々は、視覚的にシフトした観測に条件付けられた将来予測が、現在のシーンに忠実であるよりも訓練分布のコンテンツを幻覚してしまうという、繰り返し発生する現象である「訓練分布幻覚(Training-Distribution Hallucination)」を特定する。制御されたフレームトリプレット診断により、DINOv3特徴量はWan-VAE潜在変数よりも視覚的シフトに対して安定しておりながら、タスク状態の区別をより良く保持することがさらに示される。予測された将来を修正するのではなく、我々はSemantic-Temporal WAM (ST-WAM)を提案し、将来予測と履歴検索のための共有意味表現としてDINOv3を使用しつつ、細粒度のVAEダイナミクスを保持することで行動の堅牢性を向上させる。そのDual-Space Future Experts (DSFE)は将来のVAE潜在変数とDINO特徴量を共同で予測し、Current-Anchored Intent Retrieval (CAIR)は現在の視覚言語コンテキストの下で最近のDINO履歴からタスク関連の証拠を検索する。ST-WAMは、追加の身体化事前学習やタスク固有のアノテーションなしでエンドツーエンドに訓練され、推論時に明示的な将来生成を必要としない。LIBEROで98.7%、RoboTwin 2.0で92.8%を達成し、より重要なこととして、Fast-WAMと比較して、ゼロショットLIBERO-Plus性能を21.3パーセンテージポイント向上させ、視覚的シフト下での実世界成功率を25.8%から61.5%へと2倍以上に向上させる。これらの結果は、意味的時間モデリングがピクセル生成ダイナミクスを効果的に補完し、堅牢な操作を実現することを実証している。
English
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.