RynnBrain 1.1: 더욱 강력하고 일반화 가능한 체화된 기초 모델을 향하여
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
July 20, 2026
저자: Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao
cs.AI
초록
RynnBrain 1.1은 2B, 9B, 122B-A10B 규모로 구성된 임베디드 파운데이션 모델군입니다. 통합된 시공간적 및 물리적 기반 프레임워크로 학습되어 임베디드 지각, 공간 추론, 위치 파악, 계획을 지원합니다. RynnBrain 1.0과 비교하여, 이번 버전은 모델군 전반에 걸쳐 접촉점 예측을 새로 도입하고, 2B 및 9B 모델에 대해 네이티브 3D 그라운딩을 추가함으로써 로봇 조작과 더 직접적으로 정렬된 표현과 출력을 제공합니다. 또한 통합된 교차 체현 동작 공간과 체현별 마스킹을 갖춘 RynnBrain-VLA를 개발하여 Unitree G1, Astribot-S1, Tianji-Wuji에 배포했습니다. RynnBrain 1.1은 임베디드 인지, 위치 파악, 3D 그라운딩에서 뛰어난 성과를 보였으며, 특히 122B-A10B 모델은 VSI-Bench, MMSI, RefSpatial-Bench에서 평가된 모든 독점 모델 및 오픈소스 모델을 능가했습니다. 실제 로봇 실험 결과, RynnBrain으로 초기화된 정책은 Qwen 기반 및 대표적인 범용 VLA보다 우수한 성능을 보였으며, 다중 작업 및 다중 체현 공동 훈련은 작업별 훈련 대비 프로세스 점수와 성공률을 향상시켰습니다.
English
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.