ChatPaper.aiChatPaper

RoboTTT: 로봇 정책을 위한 맥락 확장

RoboTTT: Context Scaling for Robot Policies

July 16, 2026
저자: Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan
cs.AI

초록

최근 로봇 파운데이션 모델들은 단일 단계 또는 짧은 기록의 시각-운동 맥락(visuomotor context)에서 작동합니다. 본 논문에서는 추론 지연 시간 증가 없이 시각-운동 맥락을 최신 정책 대비 세 자릿수 이상 늘어난 8K 타임스텝으로 확장하는 로봇 모델 및 학습 레시피인 Test-Time-Training Robot Policies (RoboTTT)를 소개합니다. 이 맥락 길이에서 인간 영상 시연으로부터의 원샷 맥락 모방(one-shot in-context imitation), 실시간 정책 개선, 외란에 대한 강건성, 다단계 장기 과제에서의 더 강력한 성능 등 새로운 로봇 기능을 구현합니다. 또한 사전 학습 맥락 길이가 증가함에 따라 폐루프 성능이 꾸준히 향상되는 현상을 최초로 관찰합니다. RoboTTT의 핵심은 Vision-Language-Action 정책과 같은 로봇 파운데이션 모델에 Test-Time Training을 통합한 것으로, 그 결과 순환 상태가 빠른 가중치(fast weights)로 구성된 시퀀스 모델이 생성됩니다. 이 가중치들은 학습과 추론 중 모두 경사 하강법에 의해 업데이트되며, 기록을 가중치 공간으로 압축하고 장기 맥락 조건화를 위한 맥락 정보를 검색합니다. 학습 맥락 길이를 확장하기 위해 레시피는 순차 행동 강제(sequence action forcing)와 시간에 따른 절단 역전파(truncated backpropagation through time)를 결합합니다. 까다로운 실제 로봇 조작 작업에서 RoboTTT는 단일 단계 맥락 기준선 대비 전체 성능을 87% 향상시키고, 어떤 기준선도 완료하지 못했던 5분, 10단계 조립 작업을 완전히 수행합니다. 8K 타임스텝 맥락으로 학습된 RoboTTT는 동일한 모델을 1K 타임스텝으로 사전 학습한 경우보다 62% 더 높은 성능을 보여, 맥락 길이가 로봇 파운데이션 모델의 새로운 확장 축임을 시사합니다. 동영상은 https://research.nvidia.com/labs/gear/robottt/ 에서 확인할 수 있습니다.
English
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/