ChatPaper.aiChatPaper

τ_0-VLA: 세계 모델 기반 테스트 시간 연산을 활용한 계층적 로봇 기반 모델

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

August 17, 2026
저자: Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou
cs.AI

초록

장기간 로봇 조작은 로봇이 개별 기술을 안정적으로 실행하고 확장된 작업에 걸쳐 이를 일관되게 순서화하는 것을 요구한다. 대부분의 계층적 비전-언어-행동(VLA) 모델은 단일 순전파로 각각의 결정을 내리며, 어렵거나 결과에 영향을 미치는 선택에 추가 연산을 할당할 메커니즘을 제공하지 않는다. 우리는 세계 모델 기반 테스트 시 연산을 통해 고수준 하위 작업 생성을 연산 확장 가능한 추론 문제로 정식화하는 계층적 로봇 기반 모델인 τ_0-VLA를 소개한다. 각 추론 단계에서 고수준 정책은 실행 메모리를 사용하여 하위 작업을 생성하고, 필요한 경우 출력을 확정하기 전에 대안을 검색한다. 그런 다음 저수준 정책이 여러 로봇 구현체에 걸쳐 생성된 하위 작업을 실행한다. 이 정책은 다중 모달 공동 훈련을 통해 40,115시간의 이질적인 실제 데이터로 훈련된다. 도메인 내 및 분포 이동 설정에서 추가 테스트 시 연산을 할당하면 다음 하위 작업 예측 정확도가 상당히 향상되며, 이러한 이득은 장기간 로봇 조작 작업에서 더 높은 폐루프 성공으로 이어진다.
English
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.