τ_0-VLA:世界モデル誘導型テスト時計算を用いた階層的ロボット基盤モデル
τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
August 17, 2026
著者: Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou
cs.AI
要旨
長期的なロボット操作には、ロボットが個々のスキルを確実に実行するとともに、長時間にわたるタスクにおいてそれらを一貫して順序立てることが要求される。ほとんどの階層的ビジョン・言語・アクション(VLA)モデルは、このような各決定を単一のフォワードパスで行っており、困難または重要な選択に追加の計算を割り当てる仕組みを持たない。我々は、ワールドモデル誘導のテスト時計算を通じて高レベルのサブタスク生成を計算スケーラブルな推論問題として定式化する階層的ロボット基盤モデル τ_0-VLA を提案する。各推論ステップにおいて、高レベル方策は実行メモリを用いてサブタスクを生成し、必要に応じて出力を確定する前に代替案を探索する。その後、低レベル方策が生成されたサブタスクを複数のロボット身体構造にわたって実行する。本方策は、マルチモーダル共学習を用いて40,115時間の異種混合の実世界データで訓練される。ドメイン内設定と分布シフト設定の両方において、追加のテスト時計算の割り当ては次サブタスクの予測精度を大幅に向上させ、これらの利得は長期的なロボット操作タスクにおける閉ループ成功率の向上につながる。
English
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.