RxBrain:具聯合語言-視覺推理與想像的具身認知基礎模型
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
July 15, 2026
作者: Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang
cs.AI
摘要
體現認知要求智能體能將高層次的任務推理與待實現的物理狀態相連結。我們提出Hy-Embodied-RxBrain,一個具備語言與視覺聯合推理與想像能力的體現認知基礎模型。不同於強調場景理解與文字決策的視覺-語言模型,或主要預測未來視覺狀態的生成式世界模型,RxBrain在單一規劃序列中表達體現計畫,其中語言與視覺想像扮演互補角色。語言提供計畫的抽象結構,包括任務分解、規劃基元、約束條件、時間順序與決策邏輯,而視覺想像則透過世界狀態預測與聯合子目標規劃來具體化此結構,將每個規劃步驟與中間及最終的物理狀態連結起來。RxBrain採用統一的混合式Transformer多模態架構,能在單一模型中支援語言、圖像與影片的理解與生成。為訓練此能力,我們建立自動化流程,將體現影片分解為規劃步驟,並將其與視覺狀態轉換對齊,從而將影片轉換為聯合文本-視覺規劃監督信號。我們進一步提出RxBrain-Bench,以評估模型是否能透過聯合文本與視覺成分來表達體現計畫,而非僅僅進行分離的理解或生成。實驗顯示,RxBrain保有體現理解與生成能力,並能產出包含耦合文本推理、世界狀態預測與聯合子目標規劃的計畫。我們亦將RxBrain擴展至連續機器人動作生成,在不經大規模動作資料預訓練的情況下,即在真實機器人上展現出有前景的表現。這些成果為邁向體現認知基礎模型提供了初步步驟。
English
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.