BadWAM: 世界行動モデルが夢は正しいが行動は誤るとき
BadWAM: When World-Action Models Dream Right but Act Wrong
July 16, 2026
著者: Qi Li, Xingyi Yang, Xinchao Wang
cs.AI
要旨
世界行動モデル(WAM)は、身体化制御の有望な基盤として台頭しつつある。これは行動のみを予測するのではなく、行動生成と未来世界予測を結合する表現を学習する。この結合は、ロバスト性、解釈可能性、安全性の源泉とみなされることが多く、ロボットの行動が原則としてその想像上の未来と照合可能だからである。本論文では、この仮定が脆弱であることを示す。我々はBadWAMを導入する。これはWAM特有の敵対的攻撃の新たなクラスである世界行動ドリフト攻撃をモデル化・評価するための統一的枠組みであり、小さな視覚的摂動を用いてWAMが想像するものと実行するものとの間の整合性を破壊する。BadWAMは、攻撃強度とステルス性という2つの自然な基準に沿ってこの攻撃面を特徴づける。攻撃者が破壊を優先する場合、BadWAMは行動のみの敵対的攻撃を具体化し、モデルをタスク失敗行動へと直接駆動する。攻撃者がさらにステルス性を優先する場合、BadWAMは想像保存型敵対的攻撃を具体化し、モデルの予測する未来をクリーンな想像に近づけつつ、有害な行動シフトを誘発する。これら2つの攻撃は、顕在的な行動ハイジャックから、モデルがもっともらしい未来を想像しながらも非同期な行動を実行するよりステルスなケースに至るまで、WAM特有の障害のスペクトルを捉える。我々は様々なWAMの変種にわたってBadWAMを評価する。結果は、我々の攻撃が閉ループ実行下でのタスク成功率を大幅に低下させることを示す。例えば、行動のみの攻撃ではモデルの性能が成功率96.5%から43.1%に低下する。想像保存型攻撃の結果はさらに、WAM特有の脆弱性を露呈する。すなわち、適度な未来保存正則化は、未来想像ドリフトを抑えながら高い攻撃性能を維持できるということである。
English
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.