BadWAM: Wanneer wereld-actiemodellen juist dromen maar verkeerd handelen

BadWAM: When World-Action Models Dream Right but Act Wrong

July 16, 2026
Auteurs: Qi Li, Xingyi Yang, Xinchao Wang
cs.AI

Samenvatting

Wereld-actiemodellen (WAMs) komen naar voren als een veelbelovende basis voor belichaamde besturing: in plaats van alleen acties te voorspellen, leren ze representaties die actiegeneratie koppelen aan toekomstige wereldvoorspelling. Deze koppeling wordt vaak beschouwd als een bron van robuustheid, interpreteerbaarheid en veiligheid, omdat een robotactie in principe kan worden getoetst aan zijn ingebeelde toekomst. In dit artikel tonen we aan dat deze veronderstelling breekbaar is. We introduceren BadWAM, een uniform raamwerk voor het modelleren en evalueren van Wereld-Actie Drift-aanvallen: een nieuwe klasse van WAM-specifieke adversarial-aanvallen die kleine visuele verstoringen gebruiken om de afstemming tussen wat een WAM zich voorstelt en wat het uitvoert te verbreken. BadWAM karakteriseert dit aanvalsoppervlak aan de hand van twee natuurlijke criteria: aanvalssterkte en onopvallendheid. Wanneer de tegenstander prioriteit geeft aan verstoring, instantieert BadWAM een actie-only adversarial-aanval, die het model rechtstreeks naar taakfalen leidende acties drijft. Wanneer de tegenstander daarnaast prioriteit geeft aan onopvallendheid, instantieert BadWAM een verbeelding-behoudende adversarial-aanval, die schadelijke actieverschuivingen probeert te induceren terwijl de voorspelde toekomst van het model dicht bij zijn schone verbeelding blijft. Samen bestrijken deze twee aanvallen een spectrum van WAM-specifieke falen: van openlijke actiekaping tot onopvallendere gevallen waarin het model een plausibele toekomst lijkt in te beelden maar een gedesynchroniseerde actie uitvoert. We evalueren BadWAM over verschillende varianten van WAMs. Resultaten tonen aan dat onze aanvallen de taaksuccespercentages aanzienlijk verminderen onder gesloten-lus uitvoering. Zo verlaagt onze actie-only aanval de modelprestatie van 96,5% naar 43,1% succes. De resultaten van onze verbeelding-behoudende aanval leggen verder een WAM-specifieke kwetsbaarheid bloot: gematigde toekomstbehoudende regularisatie kan sterke aanvalsprestaties handhaven terwijl de afwijking in toekomstverbeelding wordt verminderd.
English
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.
PDF362July 18, 2026