LLM 예측자들이 알고 있지만 말하지 않는 것: 보정 및 충실성을 위한 내부 표현 탐구
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
July 9, 2026
저자: Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho
cs.AI
초록
예측을 위해 미세 조정된 대규모 언어 모델은 정확할 수 있지만 보정이 잘 안 될 수 있으며, 이들의 사고 사슬 추론이 예측의 근거를 충실히 반영하지 못할 수도 있다. 우리는 내부 표현이 이 두 문제에 대해 더 직접적인 창을 제공하는지 질문한다. OpenForesight 데이터셋에서 Eternis-Forecaster 8B를 사용하여 중간 활성화에 대한 표현 풀링 탐침을 훈련했으며, 이들이 훨씬 더 나은 보정을 달성함을 발견했다. 이 결과는 GLM-4.7-Flash와 GLM-4.5-Air에도 적용된다. 그런 다음 증거 제거와 오도 정보 주입을 통해 사고 사슬 충실도를 평가했다. 프롬프트에서 영향력 있는 출처를 제거하면 모델의 예측은 자주 변하지만 추론 과정은 그대로 남아 있었다. 동일한 탐침이 거짓말 탐지기로 기능했다. 이들의 활성화는 추론 과정보다 행동 변화를 훨씬 더 잘 추적했으며, 또한 사고 사슬이 교란의 영향을 숨기는 경우를 포함하여 84%의 사례에서 변화 방향을 예측했다. 마지막으로, 강제 응답은 예측이 추론이 시작되기 전에 대부분 고정됨을 드러냈다. 단일 추론 전 통과로 확정된 답변과 신뢰도가 회복되었으며, 이 미리 설정된 답변 분포의 퍼짐에 따라 질문을 라우팅하면 정확도 손실 없이 생성된 토큰을 30~47% 절약했다. 종합적으로, 이러한 결과는 내부 표현 탐침을 언어 모델 예측기 및 더 넓은 범위의 추론 모델을 보정, 감사 및 분류하기 위한 실용적인 도구로 확립한다.
English
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.