DRIFT: 敵対的パッチ攻撃によるフローマッチングVLAのデノイジング軌道の脱線
DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
August 4, 2026
著者: Hoseong Tae, Jong-Seok Lee
cs.AI
要旨
フローマッチングに基づく視覚・言語・行動(VLA)モデル、例えばpi0は、学習されたデノイジング速度場を積分することでロボットの行動を生成し、自己回帰型VLAを容易に欺く敵対的摂動に対して耐性を持つと報告されてきた。我々は、この頑健性がほとんど見せかけであることを示す。その原因は、従来の攻撃が多段階のデノイジングODEを無視していたことにある。我々は、DRIFT(Denoising Redirection via Input perturbation of the Flow-matching Trajectory、フローマッチング軌道への入力摂動によるデノイジングリダイレクション)を導入する。これは、ロボットのグリッパーに配置されるテスト時ユニバーサル敵対的パッチであり、既製のポリシーのデノイジング速度場を攻撃する。我々の中心的な発見は直感に反するものである:最初のデノイジングステップのみを攻撃することは、より広いステップ範囲を攻撃することよりも強力かつ低コストである。これは、入力空間最適化に特有の勾配の衝突によって説明され、トレーニング時のバックドア設定とは正反対である。pi0およびpi0.5において、4つのLIBEROスイートにわたって、DRIFTは小さな単一パッチで元々解けるタスクのほぼすべてを破壊し、行動空間および埋め込み空間における攻撃ベースラインをはるかに上回る。
English
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.