RoboTTT: Contextschaling voor Robotbeleid
RoboTTT: Context Scaling for Robot Policies
July 16, 2026
Auteurs: Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan
cs.AI
Samenvatting
Recente fundamentele robotmodellen werken met enkelstaps- of korte historie visuomotorische context. We introduceren Test-Time-Training Robot Policies (RoboTTT), een robotmodel en trainingsrecept dat visuomotorische context opschaalt naar 8K tijdstappen, drie ordes van grootte voorbij de state-of-the-art beleidsmodellen, zonder de inferentielatentie te verhogen. Bij deze contextlengte ontgrendelen we nieuwe robotcapaciteiten: eenmalige in-context imitatie van menselijke videodemonstraties, directe beleidsverbetering, robuustheid tegen verstoringen en betere prestaties op meerfasige langetermijntaken. We observeren ook, voor het eerst, gestage verbeteringen in closed-loop prestaties naarmate de pre-trainingscontextlengte toeneemt. In de kern integreert RoboTTT Test-Time Training in fundamentele robotmodellen zoals Visie-Taal-Actie beleid, wat resulteert in een sequentiemodel waarvan de recurrente toestand bestaat uit snelle gewichten, parameters die tijdens zowel training als inferentie door gradiëntafdaling worden bijgewerkt, waardoor geschiedenissen worden gecomprimeerd in de gewichtsruimte en contextuele informatie wordt opgehaald voor lange-contextconditionering. Om de trainingscontextlengte op te schalen, combineert het recept sequentie-actie-dwingen met getrunceerde terugpropagatie door de tijd. Bij uitdagende echte robotmanipulatietaken verbetert RoboTTT de algehele prestaties met 87% ten opzichte van de enkelstapscontext baseline en voltooit het volledig een vijf minuten durende tienfasige assemblagetaak, wat geen enkele baseline ooit doet. RoboTTT getraind met 8K tijdstappen context presteert 62% beter dan hetzelfde model getraind met 1K tijdstappen, wat suggereert dat contextlengte een nieuwe schaalas is voor fundamentele robotmodellen. Video's zijn beschikbaar op https://research.nvidia.com/labs/gear/robottt/
English
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/