Object-gecentreerde residuele RL voor zero-shot sim-naar-real VLA-verbetering
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
June 17, 2026
Auteurs: Kinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi, Jaegul Choo, Yasuyuki Matsushita
cs.AI
Samenvatting
Visie-Taal-Actiemodellen (VLA) kunnen generaliseren over uiteenlopende manipulatietaken, maar hun op imitatie leren gebaseerde beleidsstrategieën blijven kwetsbaar in precieze fysieke interacties door cumulatieve uitvoeringsfouten. Kan een reinforcement learning-beleid dat puur in simulatie is getraind de robuustheid van echte VLA’s zero-shot verbeteren? Residuele RL, dat een corrigerend beleid leert bovenop een bevroren VLA, biedt een natuurlijk raamwerk, maar bestaande benaderingen worden geconfronteerd met een fundamenteel sim-to-real-dilemma: methoden met geprivilegieerde toestanden vereisen verliesgevende distillatie voor implementatie; op beeld gebaseerde methoden lijden onder de visuele domeinkloof; en echte RL is kostbaar en onveilig. Wij stellen een objectgecentreerd residueel RL-raamwerk voor dat VLA-acties verfijnt met objectposes, wat een compacte observatieruimte mogelijk maakt die consistent overdraagt tussen simulatie en werkelijkheid. Om de twee domeinen op elkaar af te stemmen, spelen we bovendien dezelfde teleoperatiedemonstraties na in simulatie om een simulatie-tegenhanger van de echte VLA te trainen. Het residuele RL-beleid wordt uitsluitend in simulatie getraind met pose-ruisinjectie en dropout, en draagt zero-shot over naar de echte robot. Over vijf manipulatietaken op een echte Franka Research 3 (FR3)-robot verbetert onze methode het slagingspercentage van 42% naar 76% zero-shot, en de verbeterde rollouts kunnen verder worden hergebruikt om de basis-VLA opnieuw te trainen voor zelfverbetering zonder extra teleoperatie. Projectpagina: https://www.microsoft.com/en-us/research/articles/object-centric-residual-rl/
English
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions due to compounding execution errors; Can a reinforcement learning policy trained purely in simulation improve the robustness of real-world VLAs zero-shot? Residual RL, which learns a corrective policy on top of a frozen VLA, offers a natural framework, but existing approaches face a fundamental sim-to-real dilemma: privileged-state methods require lossy distillation for deployment; image-based methods suffer from the visual domain gap; and real-world RL is costly and unsafe. We propose an object-centric residual RL framework that refines VLA actions using object poses, enabling a compact observation space that transfers consistently between simulation and reality. To align the two domains, we additionally replay the same teleoperation demonstrations in simulation to train a sim counterpart of the real-world VLA. The residual RL policy is trained only in simulation with pose noise injection and dropout, and transfers zero-shot to the real robot. Across five manipulation tasks on a real Franka Research 3 (FR3) robot, our method improves the success rate from 42% to 76% zero-shot, and the improved rollouts can be further reused to retrain the base VLA for self-improvement without additional teleoperation. Project page: https://www.microsoft.com/en-us/research/articles/object-centric-residual-rl/