ChatPaper.aiChatPaper

WorldToken: 로봇 모방 학습을 위한 시간 우선 시퀀스 모델링

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

August 23, 2026
저자: Chunkai Yang, Andong Yang, Chao Gao
cs.AI

초록

로봇 정책은 각 의사결정 단계에서 이질적인 관측값을 수신하지만, 시퀀스 모델은 이러한 입력을 시간에 따라 구성하는 방식이 서로 다르다. 우리는 시간 우선 정책 구현체인 WorldToken을 소개한다. 이는 각 정책 시간 스텝 내의 다중 뷰 이미지, 고유수용감각, 과제 조건화를 하나의 월드 토큰으로 융합한다. 인과적 시간 트랜스포머가 결과적인 월드 토큰 시퀀스를 모델링하고, 확산 행동 헤드가 행동 청크를 생성한다. RoboCasa의 23개 과제에서, 동결된 사전 훈련 CLIP 텍스트 인코더를 제외하고 처음부터 훈련된 85.3M 파라미터 정책은 과제당 2,900개의 생성된 시연을 사용하여 평균 폐루프 성공률 59.45%를 달성한다. 다섯 가지 데이터셋 크기, 다섯 가지 모델 크기, 두 가지 훈련 시드에 대한 완전 요인 스윕은 추가 대상 도메인 데이터로 인한 일관된 성능 향상과 중간 크기 모델을 넘어선 지점에서의 수확 체감을 보여준다. 동일 체크포인트 이력 절단 조건에서는 가시 이력을 정책 시간 스텝 한두 개로 줄이면 50개 RoboCasa 정책 모두의 폐루프 성공률이 낮아진다. RMBench Blocks Ranking에서는 가시 이력을 146초에서 8초로 줄이면 평가자 성공률이 95%에서 28%로 낮아지는 반면, 탐색적 확장 롤아웃은 기준 스왑 시퀀스를 850초 이상 유지한다. 이러한 결과는 완전한 WorldToken 구현의 경험적 실현 가능성을 입증하고, 테스트된 레시피 하에서 데이터 스케일링 및 시간적 맥락 동작을 특성화한다. 그러나 이는 대안적인 시퀀스 구성 방식에 대한 우월성을 입증하지 않으며, 완전한 구현의 어떤 구성 요소가 관찰된 성능을 주도하는지 분리하지 않는다.
English
Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organize these inputs over time. We introduce WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token. A causal temporal Transformer models the resulting world-token sequence, and a diffusion action head generates action chunks. On 23 RoboCasa tasks, an 85.3M-parameter policy trained from scratch apart from a frozen pretrained CLIP text encoder achieves 59.45% mean closed-loop success using 2,900 generated demonstrations per task. A complete factorial sweep over five dataset sizes, five model sizes, and two training seeds shows consistent gains from additional target-domain data and diminishing returns beyond moderate model size. Under same-checkpoint history truncation, reducing visible history to one or two policy timesteps lowers closed-loop success for all 50 RoboCasa policies. On RMBench Blocks Ranking, reducing visible history from 146 to 8 seconds lowers evaluator success from 95% to 28%, while an exploratory extended rollout sustains the reference swap sequence for over 850 seconds. These results establish the empirical feasibility of the complete WorldToken instantiation and characterize its data-scaling and temporal-context behavior under the tested recipes. They do not establish superiority over alternative sequence organizations or isolate which components of the complete implementation drive the observed performance.