PhysBrain 1.5: 비전-언어 모델에서 물리 파운데이션 모델로
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
September 14, 2026
저자: DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang
cs.AI
초록
우리는 물리적 환경 이해, 행동 생성, 미래 상태 예측을 위한 통합 모델인 PhysBrain 1.5를 제안한다. 관찰, 상호작용, 환경 변화의 물리적 순환에서 동기를 얻어, 이러한 능력들을 하나의 공통 학습 프레임워크로 가져온다. 일반 비전-언어 모델에서 출발하여, 우리는 언어 응답, 엔드이펙터 움직임, 고밀도 시각 목표를 이산 시퀀스로 인코딩하고 자기회귀적 다음 토큰 예측으로 이들을 공동 최적화한다. 사전 학습은 체화 감독을 전적으로 인간 상호작용 비디오에서 얻으며, 과제 중심 에피소드를 사용하여 의미적·공간적 맥락을 복원된 움직임 및 후속 관측과 짝짓는다. 그런 다음 인간 시연, 로봇 궤적, 시뮬레이션 경험의 혼합에 대한 지도 미세 조정을 통해 모델을 적응시킨다. 28개의 체화 이해 벤치마크 전반에서 우리의 8B 모델은 평균 점수 72.5를 달성하여 새로운 오픈소스 최고 수준을 수립하고 GPT-6-Astra 및 Gemini 3.6 Flash와 같은 선도적인 독점 모델과 대등한 성능을 보인다. 이는 일반 멀티모달 능력을 유지하면서 14개 벤치마크에서 최고의 오픈소스 결과를 달성한다. 이러한 이해 평가를 넘어, 정성적 예시들은 모델이 공간적으로 정렬된 RGB, 깊이, 로봇 마스크 출력을 통해 엔드이펙터 궤적을 생성하고 미래 장면을 예측하는 능력을 보여준다.
English
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.