ChatPaper.aiChatPaper

AIレッドチーム評価が証明できることと証明できないこと

What AI Red-Team Evaluations Can and Cannot Prove

July 23, 2026
著者: Bandana Kaur
cs.AI

要旨

AIモデルのレッドチーム評価は、一部の主張を支持し、他の主張は支持しない。そして、その境界は単なる判断の問題ではなく計算可能である。我々は、評価の証拠的上限を、固定されたテスト予算の下で一つの結果が信念を変化させ得る最大の倍率として定義し、ベンチマークの帰無結果についてその上限を閉形式で導出し、それを使ってその境界を正確に特定する。我々は、計算可能な危害率を上回る場合には、適度な規模のベンチマークがあるカテゴリが明示された証拠基準を満たすことを保証し、その場合には無失敗結果が、考え得る2つの観測のうち証拠としてより強い方となり、単一の再現された失敗を上回ることを見いだす。その率を下回る場合には、実行可能な規模のいかなる受動的ベンチマークも、固定されたスコアリングルールと近似的に独立な試行構造の下で、指定された安全性の証拠を提供しない。2つの領域の間の交差は閉形式を持つ。この限界はベンチマークに限ったものではない。手順の仮説条件付き引き出し率によって表せば、これは適応的および自動化されたレッドチーミングも同様に包含し、証拠的価値を決定するのは攻撃の成功ではなく仮説間の判別であることを示す。我々は、8つの評価スイートをこの境界と照合して監査し、現在のベンチマークが高頻度の危害カテゴリには十分である一方、稀で壊滅的なカテゴリについては数桁不足していることを見いだす。安全性ベンチマークは情報がないわけではない。それらは、特定かつ計算可能な命題の集合について情報を与えるものであり、それらに求められる規律は、どの命題についてのものかを明示することである。
English
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.