ChatPaper.aiChatPaper

멈출 때를 알라: 과도한 사고를 줄이기 위한 세그먼트 수준 크레딧 할당

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

August 4, 2026
저자: Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou, Isha Slavin, Hsuan Su, Shi-Xiong Zhang, Sambit Sahu, William Campbell
cs.AI

초록

추론 언어 모델은 자주 과도하게 사고한다: 망설임, 접근 방식 포기, 자기 모순과 같이 토큰을 소비하면서도 답변을 개선하지 않는 확장된 행동 연쇄를 생성한다. 우리는 이러한 행동이 단순히 길이의 결과가 아님을 보여준다. 응답 길이를 통제하더라도, 부정확한 추론 추적은 정확한 추론 추적보다 비생산적인 자기 반성을 더 높은 비율로 나타낸다. 이를 해결하려면 자기 반성이 도움이 되는 지점과 해가 되는 지점을 식별해야 하지만, 이러한 단계별 주석을 얻는 데는 비용이 많이 든다. 우리는 추론 추적 내의 중간 답변 확정이 저렴한 대리 지표를 제공할 수 있음을 관찰한다. 추적 내의 각 최종 답변 후보를 실제 정답과 비교함으로써 추가적인 감독 없이 후속 반성이 생산적인지 여부를 결정할 수 있다. 이 통찰을 바탕으로, 우리는 각 추론 세그먼트가 정확성으로 이끄는지 또는 정확성에서 멀어지게 하는지에 따라 세그먼트 수준 신용을 할당하는 DASH(Drift Aware advantage SHaping)를 제안한다. 경쟁 수준 수학 벤치마크에서 DASH는 과도한 사고가 만연한 환경에서 가장 높은 정확도를 달성한다(평균 정확도: 59.45% vs. Dr.GRPO 58.1% vs. GRPO 56.95%). 또한 과도한 사고 행동을 줄이고 기준선보다 더 생산적인 자기 수정을 달성한다.
English
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without improving answers. We show that these behaviors are not merely a consequence of length; even when controlling for response length, incorrect traces exhibit higher rates of unproductive self-reflection than correct ones. Addressing this requires identifying where self-reflection helps vs hurts, but obtaining these step-level annotations is costly. We observe that intermediate answer commitments within reasoning traces can provide a cheap proxy: by comparing each final answer candidate in the trace to the ground truth, we can determine whether subsequent reflection is productive without any additional supervision. Building on this insight, we propose DASH (Drift Aware advantage SHaping), which assigns segment-level credit based on whether each reasoning segment leads toward or away from correctness. On competition-level math benchmarks, DASH achieves the highest accuracy where overthinking is prevalent (Average Accuracy: 59.45% vs. 58.1% Dr.GRPO vs. 56.95% GRPO) while reducing overthinking behaviors and achieving more productive self-correction than baselines.