선행 감사-수정 맥락은 LLM 검증기의 임계값을 관대한 쪽으로 이동시킨다.
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
August 17, 2026
저자: Parsa Mazaheri, Kasra Mazaheri
cs.AI
초록
자동 검사 파이프라인은 점점 더 하나의 언어 모델을 검사기로, 다른(또는 동일한) 모델을 수정기로 배치한다. 우리는 그러한 연결 구조가 검사기가 보고하는 내용을 바꾸는지 묻는다. 현재 과제를 바이트 단위로 동일하게 유지한 채 인간이 검증한 올바른 ProcessBench 트레이스에서 오경보를 측정한 결과, 완료된 감사→수리 에피소드가 이미 모델의 맥락에 있으면 15개의 모델×표현 조합 모두에서 오경보가 낮아졌다. 길이를 맞춘 비감사 대조군과 비교해 2.8~11.5퍼센트 포인트 감소했으며, 이는 해당 대조군 대비 9~25%의 감소에 해당한다. 그 방향은 누적 메시지 문헌이 예측하는 바와 모순된다. 감사가 오류를 보고한 에피소드는 오경보를 더욱 낮추었는데, 그 조작이 온전히 적용된 모델의 다섯 가지 표현 모두에서 그랬다. 부정성 비대칭이 더 많은 플래그 지정을 예측함에도 불구하고 말이다. 에피소드를 분해하면 수리 내용과 감사 판정이 상보적이다. 서로 다른 구성 요소가 서로 다른 모델군에서 그 효과를 담당한다. 신호 탐지 분석은 변화가 변별력이 아니라 역치에 있음을 보여준다. 기준(criterion)은 15개 조합 모두에서 이동했고 보정 후에도 13개에서 유지된 반면, d'는 어느 조합에서도 유지되지 않았다. 다만 d' 검정은 구성상 민감도가 절반이다. 50건의 오경보를 수동으로 점검한 결과 82%는 명백히 잘못된 것이었으므로, 이 작동 지점에서 그 이동이 반드시 해로울 필요는 없다. 추론 기능을 활성화하면 두 테스트 모델 모두에서 효과의 상대적 크기가 유지되었고, 역치 해석도 그 경우에도 성립했다.
English
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.