ChatPaper.aiChatPaper

RL^2-VLA:視覚・言語・動作モデルのためのテスト時スケーリングを備えた適応的強化学習潜在構成的ステアリング

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

July 30, 2026
著者: Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer, William Wei Jie Teo, Yuanliang Ju, Qiao Gu, Guillaume Sartoretti
cs.AI

要旨

VLA(Vision-Language-Action)モデルが実現する優れた視覚運動能力にもかかわらず、その性能は困難なタスクや領域外タスクにおいてしばしば低下する。近年の推論時ステアリングおよびスケーリング手法は、大規模なデータ収集と再訓練を必要とせずに性能を改善するが、行動サンプルはしばしば類似した振る舞いに集中したままであり、その結果、相関した失敗モードを受け継ぐことになる。さらに、既存手法は、ベースポリシーがすでに成功する可能性が高いかどうかに関係なく、すべてのタイムステップに同一の介入戦略を適用する。これらの限界に対処するため、我々はVLA潜在表現に対する強化学習を活用する適応的推論時ステアリングフレームワークであるRL^2を導入する。まず、VLA行動エキスパートから抽出した表現力豊かな潜在表現を条件とする軽量なオフライン強化学習ポリシーを訓練し、推論時にそのフロー速度を凍結されたVLAのフロー速度と合成する。この合成的ステアリング戦略は、大規模模倣学習の行動事前分布と、支配的なデモンストレーションモードを超えたオフライン強化学習によってもたらされる行動の多様性を組み合わせる。さらに、推論時ステアリングは成功状態と失敗状態で根本的に異なるスケーリング則に従うことを発見し、行動の多様性はベースVLAが失敗する可能性が高い場合に最も有益である一方、成功が確からしい場合にはすでに正確な行動を不必要に摂動させ得ることを明らかにする。この知見に基づき、RL^2は失敗が予測された場合にのみ合成的ステアリングを活性化する。SIMPLERおよびPolaRiSベンチマークにおいて、RL^2は領域外設定での成功率を最大+17.3%向上させる。一方、アブレーション研究およびスケーリング研究は、潜在表現と強化学習訓練の重要性を示す。最後に、実世界実験はこれらの利得がシミュレーションを超えて実世界へ転移することを実証し、VLA展開のための実用的かつモジュール式ステアリングフレームワークとしてRL^2を確立する。
English
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce RL^2, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, RL^2 activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, RL^2 improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing RL^2 as a practical and modular steering framework for VLA deployment.