BadWAM: 월드-액션 모델이 올바르게 꿈꾸지만 잘못 행동할 때
BadWAM: When World-Action Models Dream Right but Act Wrong
July 16, 2026
저자: Qi Li, Xingyi Yang, Xinchao Wang
cs.AI
초록
세계-행동 모델(World-action models, WAMs)은 구현된 제어를 위한 유망한 기반으로 부상하고 있다. 이 모델은 단순히 행동만 예측하는 것이 아니라, 행동 생성과 미래 세계 예측을 결합하는 표현을 학습한다. 이러한 결합은 종종 강건성, 해석 가능성, 안전성의 원천으로 간주되는데, 이는 원칙적으로 로봇의 행동이 상상된 미래와 대조하여 검증될 수 있기 때문이다. 본 논문에서는 이 가정이 취약함을 보여준다. 우리는 BadWAM을 제안하는데, 이는 세계-행동 표류 공격(World-Action Drift Attacks)을 모델링하고 평가하기 위한 통합 프레임워크이다. 이는 WAM 특유의 새로운 적대적 공격 유형으로, 작은 시각적 교란을 사용하여 WAM이 상상하는 것과 실행하는 것 사이의 정렬을 깨뜨린다. BadWAM은 이 공격 표면을 공격 강도와 은밀성이라는 두 가지 자연스러운 기준을 따라 특성화한다. 공격자가 파괴를 우선시할 때, BadWAM은 행동 전용 적대적 공격(action-only adversarial attack)을 구체화하며, 이는 모델을 과제 실패 행동으로 직접 몰아간다. 공격자가 추가로 은밀성을 우선시할 때, BadWAM은 상상 보존 적대적 공격(imagination-preserving adversarial attack)을 구체화하며, 이는 모델의 예측된 미래를 깨끗한 상상에 가깝게 유지하면서 해로운 행동 변화를 유도하려 한다. 이 두 공격은 함께 WAM 특유의 실패 스펙트럼을 포착한다. 즉, 명백한 행동 하이재킹부터 모델이 그럴듯한 미래를 상상하는 것처럼 보이지만 비동기화된 행동을 실행하는 더 은밀한 경우까지를 포함한다. 우리는 다양한 WAM 변종에 걸쳐 BadWAM을 평가한다. 결과는 우리의 공격이 폐쇄 루프 실행 하에서 과제 성공률을 상당히 감소시킴을 보여준다. 예를 들어, 행동 전용 공격은 모델 성능을 96.5% 성공률에서 43.1%로 감소시킨다. 상상 보존 공격의 결과는 WAM 특유의 취약성을 추가로 드러낸다. 즉, 적당한 미래 보존 정규화는 미래 상상 표류를 줄이면서도 강력한 공격 성능을 유지할 수 있다.
English
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.