ゼロギャップは復元ではない:層別問題別確率評価とベンチマーク汚染の段階的緩和
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
August 7, 2026
著者: Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li
cs.AI
要旨
公開ベンチマークのテストデータは必然的に事前学習コーパスへ混入し、一度記憶されると評価スコアを水増しする。汚染緩和評価は、デコード過程に介入して記憶を抑制し、汚染モデルの真の能力を回復することを目的とするが、その一般的な指標であるG-AP(Gap of Aggregate Performance、総合性能ギャップ)には欠陥がある。離散的な正誤の読み取りでは質問ごとの性能を適切に特徴づけられず、差分を取る前の平均化によって過剰抑制と過小抑制が相殺される。さらに、質問ごとに一様な重み付けを行うため、正解確率をクリーンモデルの高頻度の値へ押し上げる戦略を招く。我々はSA-PPG(Stratified Aggregate of Per-question Probability Gaps、質問ごとの確率ギャップの層別集計)を提案する。これは各質問の正解確率をサンプリングにより推定し、質問ごとにクリーンモデルとの差分を取り、クリーンモデルの正解確率に基づいて定義されたグループ内で集計するものである。既存の緩和戦略は、まず汚染箇所を推定し、その推定に対して操作を行うため、その正しさは推定の正しさに依存する。一方RailCapは生成中に汚染を判定する。サンプルが貪欲軌道に戻るたびに、軌道の次のトークンを次点に制限し、応答分布が十分に分散するまで抑制を累積する。複数の汚染モデルとベンチマークにわたる評価で、SA-PPGは既存戦略による回復が大幅に過大評価されていることを明らかにし、RailCapが最も低いSA-PPGを達成した。
English
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. Contamination mitigation evaluation intervenes in the decoding process to suppress memorization and restore a contaminated model's genuine capability, but its prevailing metric, the G-AP (Gap of Aggregate Performance), is flawed. Discrete correct/incorrect readouts cannot characterize per-question performance, averaging before differencing lets over- and under-suppression cancel out, and uniform per-question weighting invites strategies to push solve probabilities onto the clean model's high-frequency values. We propose SA-PPG (Stratified Aggregate of Per-question Probability Gaps): estimate each question's solve probability by sampling, difference it against the clean model per question, and aggregate within groups defined by the clean model's solve probability. Existing mitigation strategies first estimate where contamination lies and then operate on the estimate, so they are only as correct as the estimate. RailCap instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, accumulating suppression until the response distribution becomes sufficiently dispersed. Across multiple contaminated models and benchmarks, SA-PPG reveals that prior strategies' restoration is substantially overestimated, while RailCap attains the lowest SA-PPG.