BadWAM:当世界行动模型构想正确却行动错误
BadWAM: When World-Action Models Dream Right but Act Wrong
July 16, 2026
作者: Qi Li, Xingyi Yang, Xinchao Wang
cs.AI
摘要
世界-动作模型正成为具身控制领域一个前景广阔的基础范式:这类模型并非单纯预测动作,而是学习一种将动作生成与未来世界预测耦合的表征。这种耦合通常被视为鲁棒性、可解释性与安全性的来源,因为原则上可将机器人动作与它所想象的未来进行对比检验。本文表明,这一假设并不稳固。我们提出BadWAM这一统一框架,用于建模与评估“世界-动作漂移攻击”:这是一类针对世界-动作模型的新型对抗攻击,通过施加微小的视觉扰动,破坏模型想象内容与实际执行动作之间的对齐。BadWAM沿两个自然维度刻画这一攻击面:攻击强度与隐蔽性。当攻击者优先破坏任务时,BadWAM实例化为纯动作对抗攻击,直接驱动模型产生导致任务失败的动作;当攻击者同时强调隐蔽性时,BadWAM实例化为保想象对抗攻击,在保持模型预测未来接近干净想象的前提下,诱导有害的动作偏移。这两种攻击共同涵盖了世界-动作模型特有的失效谱系:从显式的动作劫持,到更为隐蔽的场景——模型看似想象出一个合理的未来,却执行了失同步的动作。我们在多种世界-动作模型变体上评估了BadWAM。结果表明,在闭环执行条件下,我们的攻击显著降低了任务成功率。例如,纯动作攻击将模型性能从96.5%的成功率降至43.1%。而保想象攻击的结果进一步揭示了世界-动作模型特有的脆弱性:适度保留未来的正则化项能在维持较强攻击性能的同时,减少未来想象的漂移。
English
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.