Capek 0.5:面向具身智能的執行中心視覺語言模型
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
August 7, 2026
作者: Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen
cs.AI
摘要
視覺語言模型日益成為具身智能體的推理核心。機器人執行本質上是迭代的:每個動作都會重塑場景與物理狀態,持續更新需要被感知、推理和驗證的內容。滿足這些需求需要具備互補的能力,而這些能力在監督信號、預測格式和驗證標準上各不相同。現有方法通常針對孤立的、任務特定的目標來發展這些能力,卻未闡明它們應如何圍繞整體執行進行組織與整合。我們提出 Capek 0.5,這是一個以執行為核心的能力分類體系所建構的具身視覺語言模型。該分類體系不以資料集或任務來組織訓練,而是依據具身能力在執行過程中的功能角色進行分組,包含四個能力家族:空間推理、時間理解、行動引導與狀態驗證。每個能力首先由專門的專家模型透過基於共享骨幹網路的可驗證獎勵強化學習來習得,隨後透過權重空間融合及後續的路由策略空間蒸餾,將各專家模型整合為單一的推理時模型。我們在 2B 與 35B-A3B 規模上實例化 Capek 0.5,並從三個互補的視角對其進行評估:全面的基準測試套件(包括新推出的狀態驗證基準 Capek-StateBench);從專家模型到統一模型的能力保留受控研究;以及在模擬具身環境中的閉環評估。Capek 0.5 在絕大多數對應基準測試項目上優於其初始化模型,在單一檢查點中保留了全部四種專業化能力(附帶量化損失),並能遷移至閉環具身任務執行。
English
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.