ChatPaper.aiChatPaper

τ_0-VLA:一個以世界模型引導測試時計算的階層式機器人基礎模型

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

August 17, 2026
作者: Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou
cs.AI

摘要

長期機器人操作要求機器人既能可靠地執行單一技能,也能在延伸任務中連貫地排列這些技能。大多數分層視覺-語言-動作(VLA)模型僅以單次前向傳遞做出每個此類決策,缺乏將額外計算分配給困難或重大選擇的機制。我們提出τ_0-VLA,一個分層機器人基礎模型,透過世界模型引導的測試時計算,將高階子任務生成形式化為可擴展計算的推論問題。在每個推論步驟中,高階策略使用執行記憶體生成子任務,並在必要時於提交輸出前搜尋替代方案。隨後,低階策略在多種機器人實體上執行所生成的子任務。該策略透過多模態共同訓練,在40,115小時的異質真實世界資料上進行訓練。在域內與分佈偏移的設定下,分配額外的測試時計算能大幅提升下一子任務的預測準確度,而這些提升轉化為長期機器人操作任務上更高的閉環成功率。
English
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.