ChatPaper.aiChatPaper

DRIFT:利用對抗性貼片攻擊使流匹配視覺語言動作模型的去噪軌跡脫軌

DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

August 4, 2026
作者: Hoseong Tae, Jong-Seok Lee
cs.AI

摘要

流匹配視覺-語言-行動(VLA)模型(如 pi0)透過積分學習到的去噪速度場來生成機器人動作,且據報導能抵抗輕易欺騙自迴歸 VLA 模型的對抗性擾動。我們證明這種穩健性在很大程度上是虛幻的:它源於先前攻擊忽略了多步去噪常微分方程。我們提出 DRIFT(透過流匹配軌跡的輸入擾動進行去噪重定向),這是一種置於機器人夾爪上的測試時通用對抗補丁,用以攻擊現成策略的去噪速度場。我們的核心發現違反直覺:僅攻擊第一個去噪步驟比攻擊更寬的步驟窗口更強且更便宜,我們透過輸入空間優化中獨有的梯度衝突來解釋這一現象,而該衝突恰好與訓練時後門機制相反。在跨四個 LIBERO 套件的 pi0 與 pi0.5 上,DRIFT 僅用一個小型單一補丁即可破壞幾乎所有原本可解的任務,遠遠超越動作空間與嵌入空間的攻擊基線。
English
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.