先前审计-修复情境使LLM验证器阈值向宽松偏移
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
August 17, 2026
作者: Parsa Mazaheri, Kasra Mazaheri
cs.AI
摘要
自动化检查流水线日益倾向于让一个语言模型担任检查者,另一个(或同一个)模型担任修复者。我们探究这种连接方式是否会改变检查者所报告的内容。在将当前任务保持字节完全相同的前提下,对经过人工验证正确的ProcessBench轨迹测量误报率,我们发现:模型上下文中已存在一段完整的审计→修复过程,会使15种模型×措辞组合中的全部15种组合的误报率降低,降幅为2.8至11.5个百分点(相对于长度匹配的非审计对照),即相对该对照减少9%至25%。这一方向与累积消息文献所预测的相反:在审计报告了错误的场景中,误报率进一步降低——在该操纵干净生效的模型上,全部五种措辞均如此,尽管负性不对称性预测会有更多标记。对过程进行分解后发现,修复内容与审计结论具有互补性:不同组件在不同模型家族上承载该效应。信号检测分析将该变化定位于阈值而非辨别力——判据在15/15种组合中发生移动,其中13种在校正后仍然显著,而d′在没有任何组合中保持显著,尽管d′检验在构造上灵敏度减半——对50个误报的人工审查发现其中82%确实为错误,因此在该工作点上,这一偏移未必有害。在启用推理的情况下,该效应在所测试的两个模型上均保持其相对大小,阈值解读在这两个模型上同样成立。
English
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.