事前の監査・修正コンテキストがLLM検証器の閾値を寛容さの方向へシフトさせる
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
August 17, 2026
著者: Parsa Mazaheri, Kasra Mazaheri
cs.AI
要旨
自動チェックパイプラインでは、ある言語モデルをチェッカーとして、別の(または同じ)言語モデルを修正役として配置することが増えている。我々は、そのような配線がチェッカーの報告内容を変えるかどうかを問う。人間が検証済みで正しいProcessBenchトレース上で、現在のタスクをバイト単位で同一に保ったまま誤報を測定したところ、モデルのコンテキスト内に既に完了した監査→修正のエピソードが存在すると、15/15のモデル×表現の組み合わせすべてにおいて誤報が低下することがわかった。その低下幅は、長さを一致させた非監査対照群と比較して2.8〜11.5パーセントポイントであり、同対照群に対する9〜25%の減少に相当する。この方向性は、蓄積メッセージに関する文献の予測とは矛盾する。すなわち、監査がエラーを報告したエピソードは、その操作が明確に機能したモデル上の5つの表現すべてにおいて、誤報をさらに低下させるのである。これは、ネガティビティ非対称性がより多くのフラグ付けを予測するにもかかわらずである。エピソードを分解すると、修正内容と監査判定が相補的であることがわかる。異なるコンポーネントが異なるモデルファミリー上でその効果を担っているのである。信号検出分析により、変化は識別力ではなく閾値に位置づけられる。すなわち、基準は15/15の組み合わせで移動し、13の組み合わせでは補正後も有意性を保つのに対し、d'はどの組み合わせでも補正後も有意性を保たない(ただし、d'検定は構造上、感度が半分である)。さらに、50件の誤報を手作業で監査したところ、82%が単純に誤りであり、したがってこの動作点では、そのシフトは有害である必要はない。推論を有効にした場合も、テストした両モデルで効果は相対的な大きさを維持し、閾値に関する解釈もそこでも成り立つ。
English
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.