PhysBrain 1.5:從視覺語言模型到物理基礎模型
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
September 14, 2026
作者: DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang
cs.AI
摘要
我們提出 PhysBrain 1.5,一個用於理解物理環境、生成動作與預測未來狀態的統一模型。受觀測、互動與環境變化的物理循環所啟發,我們將這些能力納入一個共同的學習框架。從通用視覺-語言模型出發,我們將語言回應、末端執行器運動與稠密視覺目標編碼為離散序列,並以自迴歸下一符元預測對其進行聯合優化。預訓練的具身監督完全來自人類互動影片,使用以任務為中心的片段,將語意與空間脈絡與還原的運動及後續觀測配對。接著,我們透過監督式微調,在人類示範、機器人軌跡與模擬經驗的混合資料上調整模型。在 28 項具身理解基準測試中,我們的 8B 模型達到平均得分 72.5,樹立新的開源最先進水準,並與 GPT-6-Astra、Gemini 3.6 Flash 等領先專有模型表現相當。它在 14 項基準測試中取得最佳開源結果,同時保留通用多模態能力。除了這些理解評估之外,定性示例顯示該模型能夠產生末端執行器軌跡,並透過空間對齊的 RGB、深度與機器人遮罩輸出預測未來場景。
English
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.