先前審計-修復情境促使LLM驗證器閾值趨向寬鬆
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
August 17, 2026
作者: Parsa Mazaheri, Kasra Mazaheri
cs.AI
摘要
自動化檢查流程越來越常將一個語言模型設置為檢查者,並將另一個(或同一個)模型設置為修復者。我們探討這種連接方式是否會改變檢查者所回報的內容。我們在當前任務保持逐位元組相同的情況下,針對經人工驗證為正確的 ProcessBench 追蹤紀錄測量誤報,結果發現:模型上下文中若已存在一段完整的「稽核→修復」過程,在全部 15 種模型×措辭組合中都會降低誤報,相對於長度匹配的非稽核對照組降低了 2.8 至 11.5 個百分點,即相較於該對照組減少 9% 至 25%。此方向與累積訊息文獻所預測的結果相反:當該過程中的稽核回報了錯誤時,誤報會進一步降低;在該操作能明確生效的模型上,五種措辭皆然,儘管負面性不對稱預期會有更多標記。將該過程分解後發現,修復內容與稽核判定具有互補性:不同組成部分在不同模型家族上發揮效果。訊號偵測分析將改變定位於閾值而非辨別力——判準在 15 種組合中全部移動,其中 13 種在校正後仍然成立,而 d' 則無一在校正後成立,儘管 d' 檢定在設計上敏感度僅為其一半。此外,對 50 個誤報的人工稽核發現,82% 純屬錯誤,因此在該操作點上,此改變未必有害。在啟用推理的情況下,該效果在所測試的兩個模型上均保持其相對大小,且閾值的解讀在該情況下同樣成立。
English
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.