ChatPaper.aiChatPaper

RxBrain: 언어-시각 공동 추론 및 상상을 갖춘 체화된 인지 기반 모델

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

July 15, 2026
저자: Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang
cs.AI

초록

체화된 인지는 에이전트가 고수준 작업 추론과 달성해야 할 물리적 상태를 연결하도록 요구한다. 우리는 언어-시각 공동 추론 및 상상력을 갖춘 체화된 인지 기반 모델인 Hy-Embodied-RxBrain을 소개한다. 장면 이해와 텍스트 기반 의사 결정을 강조하는 시각-언어 모델이나 주로 미래 시각 상태를 예측하는 생성적 세계 모델과 달리, RxBrain은 언어와 시각적 상상이 상호 보완적 역할을 하는 단일 계획 시퀀스로 체화된 계획을 표현한다. 언어는 작업 분해, 계획 기본 요소, 제약 조건, 시간적 순서 및 의사 결정 논리를 포함한 계획의 추상적 구조를 제공하는 반면, 시각적 상상은 세계 상태 예측과 공동 하위 목표 계획을 통해 이 구조를 구체화하여 각 계획 단계를 중간 및 최종 물리적 상태와 연관시킨다. RxBrain은 하나의 모델 내에서 언어, 이미지 및 비디오 이해와 생성을 지원하는 통합 다중 모드 Mixture-of-Transformers 아키텍처를 채택한다. 이러한 능력을 학습하기 위해, 우리는 비디오를 계획 단계로 분해하고 이를 시각적 상태 전환과 정렬함으로써 체화된 비디오를 공동 텍스트-시각 계획 감독으로 변환하는 자동 파이프라인을 구축한다. 또한 우리는 모델이 별도의 이해 또는 생성이 아닌 공동 텍스트 및 시각 구성 요소를 통해 체화된 계획을 표현할 수 있는지 평가하기 위해 RxBrain-Bench를 도입한다. 실험 결과는 RxBrain이 체화된 이해 및 생성 능력을 유지하며, 결합된 텍스트 추론, 세계 상태 예측 및 공동 하위 목표 계획을 포함한 계획을 생성함을 보여준다. 또한 우리는 RxBrain을 연속 로봇 행동 생성으로 확장하였으며, 대규모 행동 데이터 사전 학습 없이도 유망한 실제 로봇 성능을 보여준다. 이러한 결과는 체화된 인지를 위한 기반 모델을 향한 첫 걸음을 제공한다.
English
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.