RL²-VLA:面向視覺-語言-行動模型的自適應強化學習潛在組合式引導與測試時擴展
RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
July 30, 2026
作者: Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer, William Wei Jie Teo, Yuanliang Ju, Qiao Gu, Guillaume Sartoretti
cs.AI
摘要
儘管視覺-語言-動作(Vision-Language-Action, VLA)模型賦予了令人驚豔的視覺運動能力,但在具挑戰性與域外任務上,其表現往往會下降。近期的測試時引導與擴展方法無需大量資料收集與重新訓練即可提升效能,但動作樣本通常仍集中於相似的行為附近,因而繼承了彼此相關的失敗模式。此外,現有方法在每個時間步都施加相同的干預策略,而不論基礎策略是否已有可能成功。為了解決這些限制,我們提出 RL^2,一個自適應推論時間引導框架,利用 VLA 潛在特徵上的強化學習。首先,我們訓練一個輕量級離線強化學習策略,以從 VLA 動作專家提取的表達性潛在特徵為條件,並在推論時將其流場速度與凍結的 VLA 之流場速度進行組合。這種組合式引導策略結合了大規模模仿學習的行為先驗,以及離線強化學習在主導示範模式之外所誘發的動作多樣性。我們進一步發現,推論時間引導在成功與失敗狀態下遵循根本不同的規模定律,揭示出當基礎 VLA 可能失敗時,動作多樣性最為有益;但當成功可能性高時,動作多樣性卻可能不必要地擾動已屬準確的動作。基於此洞見,RL^2 僅在預測到失敗時才啟動組合式引導。在 SIMPLER 與 PolaRiS 基準上,RL^2 在域外設定中將成功率提升了最高 +17.3%,同時消融研究與規模研究證明了潛在表徵與強化學習訓練的重要性。最後,真實世界實驗證明這些增益可遷移至模擬之外,確立 RL^2 作為適用於 VLA 部署之實用且模組化的引導框架。
English
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce RL^2, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, RL^2 activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, RL^2 improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing RL^2 as a practical and modular steering framework for VLA deployment.