추론 잡음 제거기: 대규모 추론 모델에서 환각 탐지를 위한 추론 흔적 잡음 제거
Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models
July 24, 2026
저자: Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du
cs.AI
초록
대규모 추론 모델(LRM)은 최종 답변을 생성하기 전에 긴 추론 궤적을 생성한다. 이러한 궤적은 환각 탐지에 유용한 신호를 포함할 수 있지만, 긴 궤적에는 진실성 평가와 관련된 단서를 모호하게 만드는 잡음 단계가 자주 포함되므로 이를 활용하는 것은 간단하지 않다. 본 논문에서는 관련 없는 단계와 반복 단계라는 두 가지 일반적인 추론 잡음 형태를 식별하고, 이 두 가지 모두 환각 탐지 성능을 현저히 저하시킴을 보인다. 기존의 신뢰도 기반 점수와 단순한 임베딩 기반 필터링은 잡음이 있는 단계와 정보가 있는 단계를 안정적으로 분리하지 못한다. 이러한 문제를 해결하기 위해, 우리는 환각 탐지를 위한 추론 궤적 잡음 제거를 위한 새로운 학습 프레임워크인 REDE를 제안한다. 구체적으로, REDE는 최종 답변 주의(attention)를 자동 감독 신호로 활용하여 단계 수준 표현 공간을 형성함으로써, 잡음 단계를 안정적으로 식별하고 필터링할 수 있는 정제된 임베딩을 산출한다. REDE는 잡음 단계를 제거한 후 필터링된 추론 궤적에 적용함으로써 다양한 환각 탐지기에 쉽게 결합될 수 있다. 여러 추론 벤치마크에 대한 광범위한 실험 결과, REDE가 경쟁력 있는 기준 모델에 비해 탐지 성능을 일관되게 향상시킴을 보여준다.
English
Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include noisy steps that obscure the cues relevant to truthfulness assessment. In this paper, we identify two prevalent forms of reasoning noises, i.e., irrelevant steps and repetitive steps, and show that both substantially degrade hallucination detection performance. Existing confidence-based scores and naive embedding-based filtering fail to reliably separate noisy from informative steps. To address this challenge, we propose REDE, a novel learning framework for denoising reasoning traces for hallucination detection. Specifically, REDE leverages final-answer attention as an automatic supervision signal to shape the step-level representation space, yielding refined embeddings in which noisy steps can be reliably identified and filtered. REDE can be readily plugged into diverse hallucination detectors by operating on the filtered reasoning trajectory after removing noisy steps. Extensive experiments on multiple reasoning benchmarks show that REDE consistently improves detection performance over competitive baselines.