ChatPaper.aiChatPaper

추론 모델의 Chain-of-Thought 충실성은 선호 신호가 전달되는 위치와 방식에 따라 달라진다

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

August 29, 2026
저자: Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
cs.AI

초록

Chain-of-thought(CoT) 모니터링은 추론 트레이스가 모델의 답변을 형성하는 정보를 충실히 기록한다고 가정한다. 기존의 충실성 테스트는 명시적 편향 단서를 사용자 메시지에 배치하는 경우가 많지만, 에이전트는 도구 반환값이나 원시 산출물을 통해 선호를 접할 수 있다. 우리는 단서 위치(사용자 메시지 또는 도구 반환)와 명시성(직접 요약 또는 원시 산출물)을 변화시키는 5,100개 샘플 평가인 FACE-Eval(단서 효과의 충실한 귀인 평가)을 도입한다. 단서를 따르는 답변에서 언어화된 수용을, 단서가 포함된 모든 샘플에서 비언어화된 채택을 측정한다. 8개 계열에 속하며 총 파라미터 수가 4B에서 1.60T에 이르는 15개의 오픈 가중치 모델을 평가했다. 모든 모델에서 도구 반환 단서의 경우가 사용자 메시지 단서의 경우보다, 암묵적 단서의 경우가 명시적 단서의 경우보다 언어화된 수용이 더 낮았다. 비언어화된 채택은 15개 모델 모두에서 도구 반환 단서에 대해 더 높았고, 30개의 모델-채널 비교 중 28개에서 암묵적 단서에 대해 명시적 단서보다 더 높았다. 출처 귀인 프롬프트는 7개 모델에서 채널 격차를 좁혔으며, 때로는 사용자 채널의 비언어화된 채택을 증가시킴으로써 그 효과를 냈다. 반면, 모델에게 추론이 모니터링될 것이라고 알리는 것은 격차를 안정적으로 좁히지 못했다. 또한 각 계열의 가장 큰 모델에서 선호 채택을 탐지하기 위해 두 개의 추론 트레이스 모니터(GPT-5.6-Luna 및 GPT-4o-mini)를 사용했다. 32개의 모델-채널-명시성 셀에 걸쳐, 비언어화된 채택이 높을수록 두 모니터 모두에서 탐지 능력이 낮았다(각각 Pearson r=-0.54 및 r=-0.78). 이 결과는 선호 정보가 도구를 통해 전달되거나 원시 산출물에서 추론되어야 하는 경우, 이번 연구에서 테스트된 단일 호출, 사전 채워진 도구 설정 내에서 CoT 모니터링의 신뢰성이 낮아질 수 있음을 시사한다.
English
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.