BadWAM:當世界行動模型夢想正確卻行動錯誤
BadWAM: When World-Action Models Dream Right but Act Wrong
July 16, 2026
作者: Qi Li, Xingyi Yang, Xinchao Wang
cs.AI
摘要
世界-行動模型(WAMs)正逐漸成為具身控制領域極具前景的基礎架構:此類模型不僅預測行動,更學習將行動生成與未來世界預測相耦合的表示方式。這種耦合常被視為強健性、可解釋性與安全性的來源——因為機器人的行動原則上可經由其想像中的未來進行校驗。然而,本文揭示此假設實為脆弱。我們提出BadWAM,這是一個針對「世界-行動漂移攻擊」進行建模與評估的統一框架。此類攻擊專屬於WAM模型,利用微小視覺擾動破壞模型想像與實際執行行動之間的一致性。BadWAM依據兩個自然準則——攻擊強度與隱蔽性——刻畫此攻擊介面。當攻擊者優先考慮破壞效果時,BadWAM實例化為純行動對抗攻擊,直接驅使模型執行導致任務失敗的行動;當攻擊者同時優先考慮隱蔽性時,BadWAM則實例化為保留想像的對抗攻擊,此攻擊致力於在保持模型預測未來接近其乾淨想像的同時,誘發有害的行動偏移。兩類攻擊共同涵蓋了WAM特有的失敗光譜:從顯性的行動劫持,到更隱蔽的情況——即模型看似想像出合理的未來,卻執行失調的行動。我們在不同WAM變體上對BadWAM進行評估。結果顯示,在閉環執行條件下,我們的攻擊顯著降低了任務成功率。例如,純行動攻擊使模型成功率從96.5%降至43.1%;保留想像的攻擊結果則進一步暴露了WAM特有的脆弱性:適度的未來保留正則化雖能維持高攻擊效能,同時卻減輕未來想像的漂移。
English
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.