증폭이 예측을 의미하지는 않는다: 사고 모델의 추론 행동
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
August 13, 2026
저자: Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig
cs.AI
초록
어떤 추론 행동이 추론 모델의 정답과 연관되어 있으며, 추론 지향 훈련은 그러한 행동을 증폭하는가? 이 구분은 추론 지향 훈련이 모델 정확성과 가장 밀접하게 연결된 행동을 증폭하지 않으면서도 추론 궤적을 더 신중하게 보이게 만들 수 있기 때문에 중요하다. 우리는 행동 향상도(Behavioral Lift)라는 지표를 통해 이러한 불일치를 정량화한다. 이 지표는 모델의 추론 궤적에서 특정 행동이 존재할 때와 부재할 때 정확성이 얼마나 변화하는지를 측정한다. 텍스트 전용 및 비전-언어 추론을 포괄하는 15개 모델과 6개 벤치마크에 걸쳐, 우리는 LLM 궤적과 VLM 궤적 모두에 대해 정의된 핵심 행동을 포함하는 분류 체계로 15,282개의 추론 궤적에 주석을 달았다. 그 결과, 사고 모델이 자기 수정, 가설 검증, 불확실성 인정을 강하게 증폭하는 반면, 가장 높은 향상도를 지닌 행동은 신뢰도 교정, 지식 정렬, 자기 인식이라는 증폭-향상 격차(Amplification-Lift Gap)의 증거를 발견했다. 신뢰도 교정은 두 모달리티 모두에서 정확성의 가장 강력한 긍정적 신호 중 하나임에도 불구하고 거의 증폭되지 않으며, 불확실성 인정은 3~7배 증폭되지만 정확성과 약하게 또는 부정적으로 연관된다. 우리는 추론 지향 훈련이 가장 높은 향상도를 보이는 행동을 우선적으로 증폭하지 않음을 발견하며, 이는 표면적 형태만이 아니라 교정되고 근거가 있는 추론에 보상을 주는 프로세스 수준의 목표를 고려하게 한다.
English
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7times, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.