ChatPaper.aiChatPaper

표현 고정과 언어-행동 정렬을 통한 일반화 가능한 VLA 미세 조정

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

July 15, 2026
저자: Dwip Dalal, Shivansh Patel, Chahit Jain, Jeonghwan Kim, Utkarsh Mishra, Alex Baratian, Hyeonjeong Ha, Heng Ji, Svetlana Lazebnik, Unnat Jain
cs.AI

초록

사전 학습된 시각-언어 모델(VLM)을 로봇 시연 데이터에 행동 복제(BC) 방식으로 미세 조정하는 것은 시각-언어-행동(VLA) 정책의 표준 방법론이 되었다. 그러나 BC 미세 조정은 시각적 및 의미적 일반화를 지원하는 사전 학습된 표현을 점진적으로 덮어쓴다. 이에 대한 일반적인 해결책인 웹 이미지-텍스트 데이터 공동 학습은 이러한 문제를 방지하지 못한다. 이는 언어 손실과 행동 손실을 서로 다른 관측 데이터에 적용함으로써, 표준 조작 벤치마크가 드러내지 못하는 VLA의 언어-행동 불일치를 초래한다. 우리는 Anchor-Align을 제안한다. 이는 BC에 두 가지 목표를 추가한다. 첫째, 시각-언어 앵커링은 동결된 VLM 복사본의 계층별 표현을 증류하여 이러한 표류를 방지한다. 둘째, 언어-행동 정렬은 각 행동 목표를 이산적인 운동 방향 레이블로 변환하고, 동일한 로봇 관측 데이터에 대해 언어 예측과 행동 예측을 공동으로 학습한다. 실제 xArm7 로봇에서, 널리 사용되는 두 가지 VLA 아키텍처에 대해 Anchor-Align은 실제 로봇 성공률을 각각 28%에서 54%, 37%에서 60%로 향상시켰다. 시뮬레이션 대규모 실험에서는 LIBERO-PRO, LIBERO-Plus, CALVIN에서 각각 분포 외 변동, 지각적 강건성, 장기 제어에 걸쳐 일관된 성능 향상을 보였으며, 이는 사전 학습된 표현의 보존과 효과적인 행동 학습이 근본적으로 상충되지 않음을 시사한다. 프로젝트 페이지: anchoralignvla.github.io
English
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on web image-text data, a common remedy, does not prevent this; it applies language and action losses to separate observations, leaving VLAs with language-action misalignment that standard manipulation benchmarks do not expose. We propose Anchor-Align, which augments BC with two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy to prevent this drift, while Language-Action Alignment converts each action target into a discrete motion-direction label and jointly trains language and action prediction on the same robot observation. On a physical xArm7 robot, across two widely used VLA architectures, Anchor-Align improves real-robot success on both (28% to 54% and 37% to 60%). At scale in simulation, we demonstrate consistent improvements on OOD perturbations, perceptual robustness, and long-horizon control across LIBERO-PRO, LIBERO-Plus, and CALVIN, respectively, suggesting that preserving pretrained representations and effective action learning are not fundamentally at odds. Project page: anchoralignvla.github.io