GigaBrain-0.7:以三系統架構擴展具身基礎模型以實現湧現能力
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
August 16, 2026
作者: GigaBrain Team, Angen Ye, Axiang Sun, Can Jin, Chenxi Cheng, Chong Shi, Dengke Shang, Dingqian Zhang, Guan Huang, Guangqiang Wang, Guangqing Ding, Guo Li, Hangcong Li, Hengyu Zhong, Hongtao Lu, Jianbo Qin, Jiming Mao, Jing Zhu, Jindi Lv, Jingzhi Cui, Junjie Xie, Junyi Bao, Kai Liu, Lei Yuan, Limin Long, Lv Feng, Mingming Yu, Peng Li, Pengfei Yi, Qi Li, Qianli Zhang, Qingfang Li, Qitang Hu, Rui Zhang, Shaoyan Sun, Shibo Sun, Shiying Duan, Tenghui Chen, Tianze Liu, Weijie Ke, Wenyao Xue, Xiaofeng Wang, Xiaoyu Tian, Xinyu Liu, Xinze Chen, Yang Wang, Yankai Wang, Yejun Zeng, Yifan Li, Yifei Nie, Yilong Li, Yilong Liu, Yongchao Feng, Yumeng Wang, Yun Ye, Zhichao Liu, Ziheng He, Zonghai Yang, Zheng Zhu
cs.AI
摘要
視覺-語言-行動(VLA)模型已成為通用型具身智能體的主導範式,在結構化環境中展現出強大的複雜及長時程任務完成能力。然而,當前的VLA系統是否能從更有效的架構設計中受益、能否擴展至更大規模且更多樣化的異構數據體系,以及能否在任務與具身形態之間實現更廣泛的泛化,仍然是待解的問題。為此,我們提出了GigaBrain-0.7,一個在多種機器人具身形態上泛化能力顯著提升的具身基礎模型。具體而言,GigaBrain-0.7透過三系統架構統一了理解、預測與行動,將預訓練規模擴展至超過37,000小時的異構具身數據,並引入了單階段對齊訓練,該訓練同時優化視覺-語言理解與多具身形態行動生成。與前一代GigaBrain-0系列以及包括π_{0.5}在內的既有最先進模型相比,GigaBrain-0.7在基礎零樣本能力、語言條件指令跟隨以及後訓練任務成功率方面均取得了顯著提升。特別地,在我們自有的Maker H01平台及主流機器人具身形態上,GigaBrain-0.7在家庭與工業場景中皆展現出強大的任務適應性與完成能力。所有訓練程式碼與預訓練模型權重將全面開源釋出。
English
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.