PhysBrain 1.5:視覚言語モデルから物理基盤モデルへ
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
September 14, 2026
著者: DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang
cs.AI
要旨
我々は、物理環境を理解し、行動を生成し、将来状態を予測するための統合モデルである PhysBrain 1.5 を提案する。観測、相互作用、環境変化という物理的ループに動機づけられ、これらの能力を共通の学習枠組みに統合する。汎用視覚言語モデルを出発点として、言語応答、エンドエフェクタ運動、密な視覚ターゲットを離散系列として符号化し、自己回帰的な次トークン予測によってそれらを同時に最適化する。事前学習では、身体性に関する教師信号を完全に人間のインタラクション動画から取得し、タスク中心のエピソードを用いて、意味的文脈および空間的文脈を、復元された運動と後続の観測に対応づける。次に、人間のデモンストレーション、ロボット軌道、シミュレーション経験の混合に対して教師ありファインチューニングを行い、モデルを適応させる。28 の身体性理解ベンチマーク全体で、我々の 8B モデルは平均スコア 72.5 を達成し、オープンソースの新たな最先端を確立するとともに、GPT-6-Astra や Gemini 3.6 Flash などの主要なプロプライエタリモデルと同等の性能を示す。14 のベンチマークで最高のオープンソース結果を達成しつつ、一般的なマルチモーダル能力を維持している。これらの理解評価に加えて、定性的な例は、空間的に整合した RGB、深度、ロボットマスク出力を通じて、エンドエフェクタ軌道を生成し、将来のシーンを予測するモデルの能力を示している。
English
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.