ST-WAM:面向视觉分布偏移下鲁棒操控的语义-时间世界动作模型
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
July 31, 2026
作者: Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li
cs.AI
摘要
世界动作模型(WAMs)通过联合建模机器人动作与未来视觉动态,已成为一种前景广阔的研究范式。然而,其对像素生成式未来监督的依赖可能将动作相关的状态转移与任务无关的视觉内容纠缠在一起,从而限制了在视觉分布偏移下的鲁棒性。我们识别出训练分布幻觉这一反复出现的现象,即基于视觉偏移观测所条件化的未来预测会生成训练域内容,而非忠实于当前场景。受控的三帧诊断进一步表明,DINOv3特征在视觉偏移下保持更为稳定,同时在保留任务状态区分度方面优于Wan-VAE潜变量。我们不致力于修正预测的未来,而是提出语义-时序世界动作模型(ST-WAM),通过使用DINOv3作为共享语义表征来进行未来预测与历史检索,同时保留细粒度的VAE动态信息。其双空间未来专家(DSFE)联合预测未来VAE潜变量和DINO特征,而当前锚定意图检索(CAIR)在当前视觉-语言上下文中从近期DINO历史中检索任务相关证据。ST-WAM以端到端方式训练,无需额外的具身预训练或任务特定标注,且在推理时无需显式生成未来。它在LIBERO上达到98.7%,在RoboTwin 2.0上达到92.8%;更重要的是,与Fast-WAM相比,它将零样本LIBERO-Plus性能提升了21.3个百分点,并将视觉偏移下的真实世界成功率从25.8%提升至61.5%,提升幅度超过一倍。这些结果表明,语义-时序建模有效补充了像素生成式动态建模,从而实现鲁棒的操控。
English
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.