DRIFT: 적대적 패치 공격을 통한 플로우 매칭 VLA의 디노이징 궤적 탈선
DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
August 4, 2026
저자: Hoseong Tae, Jong-Seok Lee
cs.AI
초록
흐름 정합 비전-언어-행동(flow-matching vision-language-action, VLA) 모델(예: pi0)은 학습된 잡음 제거 속도장(denoising velocity field)을 적분하여 로봇의 행동을 생성하며, 자기회귀(autoregressive) VLA 모델을 쉽게 속이는 적대적 섭동(adversarial perturbation)에 저항하는 것으로 보고되었다. 우리는 이러한 강건성이 대체로 착각에 불과함을 보인다. 그 원인은 기존 공격들이 다단계 잡음 제거 ODE(상미분방정식)를 무시했기 때문이다. 우리는 DRIFT(Denoising Redirection via Input perturbation of the Flow-matching Trajectory)를 도입한다. 이는 테스트 시점에 로봇의 그리퍼에 부착되는 범용 적대적 패치(universal adversarial patch)로서, 사전 훈련된 정책(off-the-shelf policy)의 잡음 제거 속도장을 공격한다. 우리의 핵심 발견은 직관에 반하는 것이다. 즉, 첫 번째 잡음 제거 단계만 공격하는 것이 더 넓은 단계 구간을 공격하는 것보다 더 강력하고 계산 비용도 낮다는 것이다. 우리는 이를 입력 공간 최적화에 고유한 기울기 충돌(gradient conflict)로 설명하며, 이는 훈련 시점의 백도어(backdoor) 체계와 정확히 반대된다. pi0와 pi0.5에서 네 가지 LIBERO 벤치마크를 대상으로 한 실험에서, DRIFT는 단일 소형 패치만으로 원래 해결 가능한 거의 모든 작업을 무너뜨렸으며, 행동 공간 및 임베딩 공간 공격 기준선(action- and embedding-space attack baselines)을 훨씬 능가하는 성능을 보였다.
English
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.