ChatPaper.aiChatPaper

RL^2-VLA: 비전-언어-행동 모델을 위한 테스트-타임 스케일링 기반 적응형 RL 잠재 구성적 스티어링

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

July 30, 2026
저자: Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer, William Wei Jie Teo, Yuanliang Ju, Qiao Gu, Guillaume Sartoretti
cs.AI

초록

비전-언어-행동(Vision-Language-Action, VLA) 모델이 가능하게 하는 인상적인 시각운동 능력에도 불구하고, 어렵고 분포 외(out-of-domain) 작업에서는 성능이 종종 저하된다. 최근의 테스트 시점 스티어링 및 스케일링 방법들은 방대한 데이터 수집과 재학습 없이 성능을 개선하지만, 행동 샘플은 종종 유사한 행동 주변에 집중되어 상관된 실패 모드를 내재한다. 더욱이 기존 방법들은 기본 정책이 이미 성공할 가능성이 높은지와 무관하게 매 타임스텝마다 동일한 개입 전략을 적용한다. 이러한 한계를 해결하기 위해, 우리는 VLA 잠재 표현에 대한 강화 학습(Reinforcement Learning on VLA Latents)을 활용하는 적응형 추론 시점 스티어링 프레임워크인 RL^2를 제안한다. 첫째, 우리는 VLA 행동 전문가로부터 추출된 표현력 있는 잠재 표현에 조건화된 경량 오프라인 RL 정책을 훈련하고, 추론 중에 그 흐름 속도를 동결된 VLA의 흐름 속도와 합성한다. 이 구성적 스티어링 전략은 대규모 모방 학습의 행동 사전(behavioral prior)과 지배적인 시연 모드를 넘어서는 오프라인 RL로 유도된 행동 다양성을 결합한다. 또한 우리는 추론 시점 스티어링이 성공 상태와 실패 상태에서 근본적으로 다른 스케일링 법칙을 따른다는 것을 발견한다. 이는 행동 다양성이 기본 VLA가 실패할 가능성이 높을 때 가장 유용하지만, 성공 가능성이 높을 때는 이미 정확한 행동을 불필요하게 교란할 수 있음을 보여준다. 이 통찰력을 바탕으로 RL^2는 실패가 예측될 때만 구성적 스티어링을 활성화한다. SIMPLER 및 PolaRiS 벤치마크 전반에 걸쳐 RL^2는 분포 외 설정에서 성공률을 최대 +17.3% 향상시키며, 절제 연구와 스케일링 연구는 잠재 표현과 RL 훈련의 중요성을 입증한다. 마지막으로 실제 환경 실험은 이러한 성과가 시뮬레이션을 넘어 전이됨을 보여주며, RL^2가 VLA 배포를 위한 실용적이고 모듈식인 스티어링 프레임워크임을 확립한다.
English
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce RL^2, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, RL^2 activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, RL^2 improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing RL^2 as a practical and modular steering framework for VLA deployment.