結果が変化する場所の学習:マルチモーダル幾何のためのクレジット・アドレス可能な推論
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
August 31, 2026
著者: Jiani Guo, Junjie Wang, Jie Wu, Pengxiang Zhao, Dongdong Zhang, Shaohan Huang, Yujiu Yang, Furu Wei
cs.AI
要旨
マルチモーダル幾何推論では、VLMが精密な視覚的関係を抽出し、それを多段階の演繹的推論を通じて保持することが求められる。既存の自由形式の推論トレースは、解答を決定づける判断を不明瞭にし、軌跡レベルの強化学習は、単一の終端信号を応答全体に分配する。本稿では、クレジット割り当て可能な推論を導入する。これは、推論中に顕在化する意味単位が、学習において代替案の比較とクレジット割り当てが行われる場所をも定義するという原理に基づく。この原理を、図を保持しつつ視覚的関係を行アドレス指定可能な実行可能コードとして表現し、推論を型付きイベントに組織化するCode-CoTと、構造的事前分布と型正規化エントロピーを用いてイベント境界を選択し、共有プレフィックスから完全な継続をサンプリングして、結果の差異を局所的なアドバンテージに変換するCE-GRPOという二つの手法によって具現化する。9つの幾何学ベンチマークにおいて、CE-GRPOは平均精度76.04を達成し、Qwen3-VL-8Bおよび軌跡レベルGRPOをそれぞれ8.09ポイントおよび3.43ポイント上回る。その相対的優位性は中間イベント数の増加に伴って拡大し、長く依存関係の強いマルチモーダル推論における表現と最適化の共設計の価値を実証している。
English
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-level reinforcement learning distributes a single terminal signal across the entire response. We introduce credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit. We instantiate this principle with Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events, and CE-GRPO, which selects event boundaries using structural priors and type-normalized entropy, samples complete continuations from shared prefixes, and converts outcome differences into localized advantages. Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by 8.09 and 3.43 points, respectively. Its relative advantage increases with the number of intermediate events, demonstrating the value of representation--optimization co-design for long, dependency-heavy multimodal reasoning.