増幅は予測的を意味しない:思考モデルにおける推論行動
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
August 13, 2026
著者: Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig
cs.AI
要旨
推論モデルにおいて、どの推論行動が正答と関連しているのか。また、推論指向の訓練はそうした行動を増幅するのか。この区別は重要である。なぜなら、推論指向の訓練は、モデルの正答性と最も結びついた行動を増幅することなく、トレースをより慎重に検討したように見せかけうるからである。我々はこの不整合を、モデルの推論トレースにおいてある行動が存在する場合と存在しない場合とで正答率がどれだけ変化するかを測定する指標であるBehavioral Lift(行動リフト)によって定量化する。テキストのみおよび視覚言語推論を対象とする6ベンチマーク、15モデルにわたり、我々はLLMトレースとVLMトレースの両方に定義された中核行動からなる分類体系を用いて15,282件のトレースを注釈付けした。その結果、思考モデルが自己修正、仮説検証、不確実性の表明を強く増幅する一方、最も高いリフトを示す行動は信頼度較正、知識整合性、自己認識であるという、増幅-リフト乖離(Amplification-Lift Gap)の証拠を得た。信頼度較正は両モダリティにおいて正答の最も強い正のシグナルの一つであるにもかかわらず、ほとんど増幅されない。不確実性の表明は3〜7倍に増幅されるにもかかわらず、正答との関連は弱いか負である。推論指向の訓練は最もリフトの高い行動を優先的に増幅するわけではないことが分かり、表面の形式だけでなく、較正され根拠に基づいた推論に報酬を与えるプロセスレベルの目的関数が動機づけられる。
English
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7times, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.