ChatPaper.aiChatPaper

GigaBrain-0.7: 3-시스템 아키텍처를 통한 구현형 파운데이션 모델의 창발적 능력으로의 확장

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

August 16, 2026
저자: GigaBrain Team, Angen Ye, Axiang Sun, Can Jin, Chenxi Cheng, Chong Shi, Dengke Shang, Dingqian Zhang, Guan Huang, Guangqiang Wang, Guangqing Ding, Guo Li, Hangcong Li, Hengyu Zhong, Hongtao Lu, Jianbo Qin, Jiming Mao, Jing Zhu, Jindi Lv, Jingzhi Cui, Junjie Xie, Junyi Bao, Kai Liu, Lei Yuan, Limin Long, Lv Feng, Mingming Yu, Peng Li, Pengfei Yi, Qi Li, Qianli Zhang, Qingfang Li, Qitang Hu, Rui Zhang, Shaoyan Sun, Shibo Sun, Shiying Duan, Tenghui Chen, Tianze Liu, Weijie Ke, Wenyao Xue, Xiaofeng Wang, Xiaoyu Tian, Xinyu Liu, Xinze Chen, Yang Wang, Yankai Wang, Yejun Zeng, Yifan Li, Yifei Nie, Yilong Li, Yilong Liu, Yongchao Feng, Yumeng Wang, Yun Ye, Zhichao Liu, Ziheng He, Zonghai Yang, Zheng Zhu
cs.AI

초록

비전-언어-행동(VLA) 모델은 범용 체화 에이전트의 지배적인 패러다임으로 부상하여, 구조화된 환경에서 복잡하고 장기 지평의 과제 수행에 강력한 성능을 입증해 왔다. 그러나 현재의 VLA 시스템이 보다 효과적인 아키텍처 설계의 이점을 얻을 수 있는지, 훨씬 더 크고 이질적인 데이터 체계로 확장될 수 있는지, 그리고 과제와 구현체 전반에 걸쳐 더 폭넓은 일반화를 달성할 수 있는지는 여전히 미해결 문제로 남아 있다. 이러한 문제를 해결하기 위해, 우리는 다양한 로봇 구현체에 걸쳐 현저히 개선된 일반화 성능을 갖춘 체화 기반 모델인 GigaBrain-0.7을 제시한다. 구체적으로, GigaBrain-0.7은 3-시스템 아키텍처를 통해 이해, 예측, 행동을 통합하고, 37,000시간 이상의 이질적 체화 데이터로 사전학습을 확장하며, 비전-언어 이해와 다중 구현체 행동 생성을 동시에 최적화하는 단일 단계 정렬 훈련을 도입한다. 이전 GigaBrain-0 시리즈 및 π_{0.5}를 포함한 기존 최첨단 모델과 비교하여, GigaBrain-0.7은 기반 제로샷 성능, 언어 조건부 지시 수행, 후속 훈련 과제 성공률에서 상당한 개선을 달성한다. 특히, 자체 개발한 Maker H01 플랫폼과 주류 로봇 구현체에서 GigaBrain-0.7은 가정 및 산업 현장 시나리오 모두에서 강력한 과제 적응성과 완료 능력을 입증한다. 모든 훈련 코드와 사전학습 모델 가중치는 공개될 예정이다.
English
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.