文書抽出のための妥当なフィールド単位の選択的リスク制御:三つの失敗モード、妥当性ラダー、そして条件付けが有効な場合
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
July 28, 2026
著者: Bhaskar Gurram
cs.AI
要旨
フィールド単位の受理・レビューにおいて、選択的リスクを最大αに抑えること——受理されたフィールドのみの誤り率を管理すること——は、文書抽出システムが要求する信頼契約であり、標準的な手続きは実際の文書上でこれを黙って違反する。800件のCORDレシートから得られた13,859件の実データclaude-sonnet-5フィールド(正解率49.0%)において、我々は3つの失敗モードを診断する: 文書クラスタリング(デザイン効果1.84〜2.45)、スコア再適合リーク(リスク0.127でカバレッジ0.416、α=0.10を95%の分割で違反)、およびタイマス病理(縮退スコアが閾値グリッドを崩壊させ、0.030から0.001へ)。我々は修正策を妥当性ラダーとして整理し、各層ごとに保証形式を明示する。適合/検証分割プロトコルは、学習された融合に対して期待選択リスク制御を回復する: 名目α=0.10でカバレッジ0.318、リスク0.096、許容帯なし(本番バリアントは0.326)——これは平均的な点であり、再分割の47.5%では実現リスクがαを上回る。証明書ではない。Mondrian Learn-then-Testと正確な二項分布の裾を用いることで、グループごとのPAC証明書が得られる: フィールドiidでは0.171(リスク0.068)、クラスタ補正では0.140、文書iidでは0.060——文書に整合する唯一の層であり、正直に言って現在はほぼ空虚である。事前指定された来歴分類法であるサポートビンは、sonnet CORD収集データ上のすべての厳密性層で勝利する(p<1e-4、Bonferroni補正後)——しかしこの勝利は、haikuまたはqwenでは同じ文書上で再現しない——一方、高精度コーパスではプールされた閾値が勝利する: 条件付けは、プール法が証明できないまさにその状況で役立ち、他の状況では学習スコアに包含される。選択プロセスに未使用のclaude-haiku-4-5における凍結構成での確認は、両方のリスク水準で維持された。また、盲検の3名のアノテータによる人間ゴールド監査は、実用層の受理集合リスクが10%の予算に対して1.3%であることを検証する(Fleissのκ=0.83; ラベルは一方的に悲観的な方向に誤る)。Apache-2.0で、シード固定・回帰ゲート付き手続きとともに公開されている。
English
Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- is the trust contract document-extraction systems need, and the natural procedure silently violates it on real documents. On 13,859 genuine claude-sonnet-5 fields from 800 CORD receipts (49.0% correct) we diagnose three failure modes: document clustering (design effect 1.84-2.45), score-refit leakage (coverage 0.416 at risk 0.127, violating alpha=0.10 in 95% of splits), and a tie-mass pathology (a degenerate score collapses the threshold grid, 0.030 to 0.001). We organize the fixes as a validity ladder, guarantee form stated per tier. A fit/val split protocol restores expected-selective-risk control for a learned fusion: coverage 0.318 at risk 0.096 at nominal alpha=0.10, no tolerance band (production variant 0.326) -- an on-average point whose realized risk exceeds alpha in 47.5% of resplits, not a certificate. Mondrian Learn-then-Test with exact binomial tails yields per-group PAC certificates: field-iid 0.171 at risk 0.068, cluster-corrected 0.140, doc-iid 0.060 -- the only tier matching documents, honestly near-vacuous today. Support-bin, the pre-specified provenance taxonomy, wins every rigor tier on the sonnet CORD capture (p<1e-4, Bonferroni-corrected) -- a win that does not replicate on the same documents under haiku or qwen -- while on higher-accuracy corpora pooled thresholds win: conditioning helps exactly where pooled cannot certify, subsumed by a learned score elsewhere. A frozen-configuration confirmation on selection-untouched claude-haiku-4-5 held at both risk levels, and a blind three-annotator human-gold audit verifies the practical tier's accepted-set risk at 1.3% against its 10% budget (Fleiss' kappa=0.83; labels err one-sidedly pessimistic). Released Apache-2.0 with seed-pinned, regression-gated procedures.