Capek 0.5:面向具身智能的以执行为中心的视觉语言模型
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
August 7, 2026
作者: Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen
cs.AI
摘要
视觉-语言模型正日益成为具身智能体的推理核心。机器人执行本质上具有迭代性:每个动作都会重塑场景和物理状态,持续更新需要被感知、推理和验证的内容。满足这些需求需要具备在监督信号、预测格式和验证标准上各不相同的互补能力。现有方法通常针对孤立的、任务特定的目标来发展这些能力,而未阐明它们应如何围绕执行整体进行组织和整合。我们提出Capek 0.5,一个以执行为中心的能力分类体系为基础的具身视觉-语言模型。该分类体系不按数据集或任务组织训练,而是根据具身能力在执行过程中的功能角色对其进行分组,包含四个能力族:空间推理、时间理解、动作引导和状态验证。每种能力首先由专属专家模型通过基于共享主干网络的可验证奖励强化学习获得,随后通过权重空间合并及路由策略空间蒸馏,将各专家模型整合为单一的推理时模型。我们在2B和35B-A3B规模上实例化Capek 0.5,并从三个互补视角进行评估:综合基准测试套件(包括Capek-StateBench——一个新的状态验证基准);从专家模型到统一模型的能力保持受控研究;以及在模拟具身环境中的闭环评估。Capek 0.5在大多数匹配的基准测试条目上相对于其初始化模型有所提升,在一个检查点中以可量化的损失保留了全部四种专项能力,并成功迁移到闭环具身任务执行中。
English
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.