ChatPaper.aiChatPaper

增强并不意味着可预测:思考模型中的推理行为

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

August 13, 2026
作者: Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig
cs.AI

摘要

在推理模型中,哪些推理行为与正确答案相关联?以推理为导向的训练是否放大了这些行为?这一区分之所以重要,是因为以推理为导向的训练可以使推理轨迹看起来更具深思熟虑的特征,却并未放大与模型正确性最紧密关联的那些行为。我们通过行为提升(Behavioral Lift)这一指标来量化这种不匹配,该指标衡量某一行为在模型推理轨迹中出现与不出现时正确性变化的幅度。横跨15个模型和6个涵盖纯文本推理与视觉-语言推理的基准测试,我们以一套核心行为同时适用于大语言模型(LLM)和视觉语言模型(VLM)轨迹的分类体系,对15,282条轨迹进行了标注。我们发现了放大-提升差距(Amplification-Lift Gap)的存在证据:思维模型强烈放大自我修正、假设检验和不确定性承认,而提升值最高的行为则是置信度校准、知识对齐和自我意识。置信度校正在两种模态中都是正确性最强的正向信号之一,却几乎没有被放大;不确定性承认被放大了3至7倍,却与正确性呈弱相关或负相关。我们发现,以推理为导向的训练并未优先放大提升值最高的行为,这凸显了引入过程层面目标函数的必要性,以奖励经过校准且具有依据的推理,而非仅关注表面形式。
English
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7times, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.