ChatPaper.aiChatPaper

RynnBrain 1.1: より高性能で汎用的な身体化基盤モデルを目指して

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

July 20, 2026
著者: Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao
cs.AI

要旨

本稿では、RynnBrain 1.1ファミリーを発表する。これは2B、9B、122B-A10Bという規模にわたる身体化基盤モデル群である。統一された時空間的かつ物理的に接地されたフレームワークで学習され、身体化知覚、空間推論、位置特定、計画をサポートする。RynnBrain 1.0と比較して、本ファミリーでは全モデルにおける接触点予測、および2Bと9Bモデルにおけるネイティブな3Dグラウンディングを新たに導入し、ロボット操作とより直接的に整合する表現と出力を実現する。また、統一されたクロス身体化行動空間と身体化固有のマスキングを備えたRynnBrain-VLAを開発し、Unitree G1、Astribot-S1、Tianji-Wujiに展開した。RynnBrain 1.1は、身体化認知、位置特定、3Dグラウンディングにおいて強力な結果を示し、122B-A10BモデルはVSI-Bench、MMSI、RefSpatial-Benchにおいて評価されたすべてのプロプライエタリおよびオープンソースモデルを上回った。実機ロボット実験では、RynnBrainで初期化されたポリシーがQwenベースおよび代表的な汎用VLAを凌駕し、さらに、マルチタスク・マルチ身体化統合訓練が、タスク個別訓練と比較してプロセススコアと成功率を向上させることを示した。
English
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.