放大不等於可預測:思考模型中的推理行為
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
August 13, 2026
作者: Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig
cs.AI
摘要
在推理模型中,哪些推理行為與正確答案相關?推理導向訓練是否會放大這些行為?此區別之所以重要,是因為推理導向訓練可能使軌跡看起來更具深思熟慮性,卻未必放大與模型正確性最相關的行為。我們透過行為提升(Behavioral Lift)來量化此不匹配,該指標衡量在模型的推理軌跡中,某行為出現與否時正確性的變化程度。在涵蓋純文字與視覺-語言推理的15個模型與6個基準測試中,我們以一套核心行為同時為LLM與VLM軌跡定義的分類法,標註了15,282條軌跡。我們發現了放大-提升差距(Amplification-Lift Gap)的證據,其中思考模型大幅放大自我修正、假設檢驗與不確定性承認,而提升最高的行為則是信心校準、知識對齊與自我覺察。信心校準在兩種模態中皆是最強的正確性正向訊號之一,卻幾乎未被放大;不確定性承認被放大了3至7倍,卻與正確性僅有弱相關或負相關。我們發現推理導向訓練並未優先放大提升最高的行為,這促使我們採用能夠獎勵校準且紮根之推理的過程層級目標,而非僅獎勵表面形式。
English
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7times, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.