DRIFT:利用对抗补丁攻击偏离流匹配视觉-语言-动作模型的去噪轨迹
DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
August 4, 2026
作者: Hoseong Tae, Jong-Seok Lee
cs.AI
摘要
诸如 pi0 之类的流匹配视觉-语言-动作(VLA)模型通过积分学习到的去噪速度场来生成机器人动作,并且据报道能够抵抗那些容易欺骗自回归 VLA 的对抗扰动。我们表明,这种鲁棒性在很大程度上是一种错觉:它源于先前的攻击忽略了多步去噪常微分方程(ODE)。我们提出了 DRIFT(通过对流匹配轨迹的输入扰动进行去噪重定向),这是一种放置在机器人夹爪上的测试时通用对抗补丁,用于攻击现成策略的去噪速度场。我们的核心发现是反直觉的:仅攻击去噪的第一步比攻击更宽的时间步窗口既更强又成本更低,我们通过输入空间优化特有的梯度冲突来解释这一点,而这与训练时后门机制恰好相反。在四个 LIBERO 套件上的 pi0 和 pi0.5 中,DRIFT 用一个很小的单一补丁便基本攻破了所有原本可解决的任务,远超动作空间和嵌入空间的攻击基线。
English
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.