ChatPaper.aiChatPaper

Capek 0.5: 具現化知能のための実行中心の視覚言語モデル

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

August 7, 2026
著者: Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen
cs.AI

要旨

視覚言語モデルは、身体化エージェントの推論の中核としてますます重要な役割を果たしつつある。ロボットの実行は本質的に反復的であり、各動作がシーンと物理状態を再形成し、知覚され、推論され、検証されるべきものを絶えず更新する。こうした要求に応えるには、教師信号、予測形式、検証基準が異なる相補的な能力が必要となる。既存の手法は通常、これらの能力を孤立したタスク固有の目的に対して開発しており、実行全体を中心にそれらをどのように組織化し統合するかは未解決のままである。本稿では、実行中心の能力分類法に基づいて構築された身体化視覚言語モデルCapek 0.5を提案する。この分類法は、データセットやタスクごとに訓練を構成するのではなく、実行全体を通じた機能的役割に従って身体化能力をグループ化し、空間推論、時間理解、行動指導、状態検証の4つの能力群から構成される。各能力はまず、共有バックボーンからの検証可能な報酬を用いた強化学習によって専用のスペシャリストが獲得し、その後、スペシャリストは重み空間マージとそれに続くルーティング付きポリシー空間蒸留によって、単一の推論時モデルに統合される。我々はCapek 0.5を2Bおよび35B-A3Bスケールで実装し、状態検証のための新しいベンチマークであるCapek-StateBenchを含む包括的なベンチマークスイート、スペシャリストから統合モデルへの能力保持に関する統制研究、シミュレーションされた身体化環境におけるクローズドループ評価という3つの相補的な観点から評価する。Capek 0.5は、対応するベンチマーク行の大部分で初期化時よりも改善し、4つすべての専門能力を定量化された損失とともに単一のチェックポイントに保持し、クローズドループの身体化タスク実行へ転移する。
English
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.