ChatPaper.aiChatPaper

RL²-VLA: 面向视觉-语言-动作模型的自适应强化学习潜在组合引导与测试时扩展

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

July 30, 2026
作者: Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer, William Wei Jie Teo, Yuanliang Ju, Qiao Gu, Guillaume Sartoretti
cs.AI

摘要

尽管视觉-语言-动作(VLA)模型赋予了令人印象深刻的视觉运动能力,但它们在具有挑战性和域外任务上的性能常常会下降。近期的测试时调控和缩放方法无需大量数据收集和重新训练即可提升性能,但动作样本往往仍集中在相似行为附近,因此继承了相关的失败模式。此外,现有方法在每个时间步都采用相同的干预策略,而不考虑基础策略是否已经可能成功。为了解决这些局限性,我们提出了RL^2,一种利用VLA潜变量上的强化学习的自适应推理时调控框架。首先,我们训练一个轻量级离线强化学习策略,该策略以从VLA动作专家中提取的表达性潜变量为条件,并在推理时将其流速度与冻结的VLA的流速度进行组合。这种组合式调控策略将大规模模仿学习的行为先验与离线强化学习在主导示范模式之外引发的动作多样性相结合。我们进一步发现,推理时调控在成功和失败状态下遵循根本不同的缩放定律,这表明当基础VLA可能失败时,动作多样性最为有益,但当成功可能时,它会不必要地扰动已经准确的动作。基于这一见解,RL^2仅在预测到失败时才激活组合式调控。在SIMPLER和PolaRiS基准测试中,RL^2在域外设置中将成功率提升高达17.3%,同时消融实验和缩放研究证明了潜变量表示和强化学习训练的重要性。最后,真实世界实验表明,这些收益能够迁移到模拟之外,使RL^2成为面向VLA部署的实用且模块化的调控框架。
English
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce RL^2, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, RL^2 activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, RL^2 improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing RL^2 as a practical and modular steering framework for VLA deployment.