PhysBrain 1.5:从视觉-语言模型到物理基础模型
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
September 14, 2026
作者: DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang
cs.AI
摘要
我们提出 PhysBrain 1.5,这是一个用于理解物理环境、生成动作和预测未来状态的统一模型。受观察、交互和环境变化构成的物理闭环启发,我们将这些能力纳入一个共同的学习框架。从通用视觉-语言模型出发,我们将语言响应、末端执行器运动和稠密视觉目标编码为离散序列,并通过自回归下一词元预测对其进行联合优化。预训练完全从人类交互视频中获取具身监督,利用以任务为中心的片段将语义和空间上下文与恢复的运动及随后的观测配对。随后,我们通过在人类演示、机器人轨迹和模拟经验的混合数据上进行监督微调来适配模型。在 28 个具身理解基准上,我们的 8B 模型取得了 72.5 的平均分,创下新的开源最先进水平,并与 GPT-6-Astra 和 Gemini 3.6 Flash 等领先专有模型表现相当。它在 14 个基准上取得了最佳开源结果,同时保持了通用多模态能力。除这些理解评估外,定性示例展示了该模型生成末端执行器轨迹,以及通过空间对齐的 RGB、深度和机器人掩码输出预测未来场景的能力。
English
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.