ChatPaper.aiChatPaper

τ_0-VLA:一种具有世界模型引导的测试时计算的层次化机器人基础模型

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

August 17, 2026
作者: Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou
cs.AI

摘要

长时程机器人操作要求机器人既能可靠地执行单个技能,又能在扩展任务中对技能进行连贯排序。大多数分层视觉-语言-动作(VLA)模型通过单次前向传播完成每一次决策,缺乏为困难或关键决策分配额外计算的机制。我们提出τ_0-VLA,一种分层机器人基础模型,将高层子任务生成表述为一种通过世界模型引导的测试时计算实现计算可扩展的推理问题。在每个推理步骤中,高层策略利用执行记忆生成子任务,并在必要时先对备选方案进行搜索再确定输出。随后,底层策略在多种机器人形态上执行生成的子任务。该策略在40,115小时异构真实世界数据上训练,并采用多模态协同训练。在域内与分布偏移场景下,分配额外的测试时计算能显著提升下一子任务预测准确率,且这些收益能够转化为长时程机器人操作任务上更高的闭环成功率。
English
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.