ST-WAM: 시각적 분포 변화 하에서의 강건한 조작을 위한 의미론적-시간적 세계 행동 모델
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
July 31, 2026
저자: Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li
cs.AI
초록
월드 액션 모델(WAM)은 로봇 행동과 미래 시각적 역학을 공동으로 모델링하는 유망한 패러다임으로 부상하였다. 그러나 이들은 픽셀 생성적 미래 감독에 의존함으로써, 행동 관련 상태 전이와 작업 무관 시각적 콘텐츠가 얽히게 되어 시각적 분포 변화 하에서 강건성이 제한될 수 있다. 우리는 훈련-분포 환각(Training-Distribution Hallucination)을 식별하였는데, 이는 시각적으로 이동된 관측값에 조건화된 미래가 현재 장면에 충실하게 유지되지 못하고 훈련 도메인 콘텐츠를 환각하는 반복적 현상이다. 통제된 프레임-삼중쌍 진단은 DINOv3 특징이 Wan-VAE 잠재 변수보다 시각적 변화 전반에 걸쳐 더 안정적으로 유지되면서 작업 상태 구분을 더 잘 보존함을 추가로 보여준다. 우리는 예측된 미래를 교정하는 대신, 의미-시간적 WAM(ST-WAM)을 제안하여 DINOv3를 미래 예측 및 이력 검색을 위한 공유 의미 표현으로 사용하면서 세밀한 VAE 역학을 유지함으로써 행동 강건성을 개선한다. 이중 공간 미래 전문가(DSFE)는 미래 VAE 잠재 변수와 DINO 특징을 공동으로 예측하고, 현재 고정 의도 검색(CAIR)은 현재 시각-언어 맥락 하에서 최근 DINO 이력으로부터 작업 관련 증거를 검색한다. ST-WAM은 추가적인 체화 사전 훈련이나 작업별 주석 없이 종단 간 훈련되며, 추론 시 명시적 미래 생성을 요구하지 않는다. 이 모델은 LIBERO에서 98.7%, RoboTwin 2.0에서 92.8%를 달성한다. 더 중요하게는, Fast-WAM과 비교하여 제로샷 LIBERO-Plus 성능을 21.3퍼센트 포인트 향상시키고, 시각적 변화 하에서 실제 로봇 성공률을 25.8%에서 61.5%로 두 배 이상 증가시킨다. 이러한 결과는 의미-시간적 모델링이 강건한 조작을 위해 픽셀 생성적 역학을 효과적으로 보완함을 입증한다.
English
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. We identify Training-Distribution Hallucination, a recurring phenomenon in which futures conditioned on visually shifted observations hallucinate training-domain content rather than remain faithful to the current scene. A controlled frame-triplet diagnosis further shows that DINOv3 features remain more stable across visual shifts while better preserving task-state distinctions than Wan-VAE latents. Rather than correcting the predicted futures, we propose Semantic-Temporal WAM (ST-WAM) to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics. Its Dual-Space Future Experts (DSFE) jointly predict future VAE latents and DINO features, while Current-Anchored Intent Retrieval (CAIR) retrieves task-relevant evidence from recent DINO history under the current visual-language context. ST-WAM is trained end-to-end without additional embodied pretraining or task-specific annotations, and requires no explicit future generation at inference. It achieves 98.7% on LIBERO and 92.8% on RoboTwin 2.0; more importantly, compared with Fast-WAM, it improves zero-shot LIBERO-Plus performance by 21.3 percentage points and more than doubles real-world success under visual shifts from 25.8% to 61.5%. These results demonstrate that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.