RxBrain: 言語と視覚の統合的推論と想像を備えた身体化認知基盤モデル
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
July 15, 2026
著者: Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang
cs.AI
要旨
身体化認知は、エージェントが高レベルのタスク推論と達成すべき物理的状態を結びつけることを必要とする。我々は、言語‐視覚の共同推論と想像力を備えた身体化認知基盤モデルHy-Embodied-RxBrainを紹介する。シーン理解やテキストによる意思決定を重視する視覚言語モデルや、主に未来の視覚状態を予測する生成的世界モデルとは異なり、RxBrainは単一の計画シーケンス内で身体化計画を表現し、言語と視覚想像力が補完的な役割を果たす。言語はタスク分解、計画プリミティブ、制約、時間的順序、意思決定ロジックなど計画の抽象的な構造を提供し、視覚想像力は世界状態予測と共同サブゴール計画を通じてこの構造を具体化し、各計画ステップを中間および最終的な物理的状態に関連付ける。RxBrainは統一されたマルチモーダルなMixture-of-Transformersアーキテクチャを採用し、言語、画像、動画の理解と生成を単一モデル内で実現する。この能力を訓練するために、我々は身体化ビデオを計画ステップに分解し、それらを視覚状態遷移と整合させることで、ビデオをテキスト‐視覚共同計画の教師信号に変換する自動パイプラインを構築する。さらに、モデルが個別の理解や生成ではなく、テキストと視覚の共同コンポーネントを通じて身体化計画を表現できるかを評価するために、RxBrain-Benchを導入する。実験により、RxBrainは身体化理解と生成の能力を維持し、結合されたテキスト推論、世界状態予測、共同サブゴール計画を伴う計画を生成することが示された。また、RxBrainを連続的なロボット動作生成に拡張し、大規模な動作データによる事前学習なしで有望な実ロボット性能を示す。これらの結果は、身体化認知のための基盤モデルに向けた最初の一歩を提供する。
English
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.