ChatPaper.aiChatPaper

AI紅隊評測能證明什麼、不能證明什麼

What AI Red-Team Evaluations Can and Cannot Prove

July 23, 2026
作者: Bandana Kaur
cs.AI

摘要

對AI模型的紅隊評估,會支持某些主張,而不支持其他主張;而這兩者之間的界線是可以計算的,而不只是判斷的問題。我們將一項評估的證據上限定義為:在固定測試預算下,某個結果能改變信念的最大倍數;並針對基準測試的零結果,以閉式形式推導出該上限,再用它精確定位這條界線。我們發現,當危害率高於某個可計算數值時,規模適中的基準測試即可依既定的證據標準認證某個類別;此時,零失敗紀錄是兩種可能觀察結果中較強的一種,其分量勝過單一的可重現失敗。當危害率低於該數值時,在固定評分規則與近似獨立的試驗結構下,沒有任何規模可行的被動基準測試能提供所指定的安全性證據。這兩個機制之間的交叉點具有閉式解。該上限並非基準測試所獨有;以某個程序在假設條件下的引出率來表述,它同樣涵蓋自適應與自動化的紅隊測試,並顯示出,決定證據價值的是假設之間的區辨,而非攻擊是否成功。我們對照這條界線審計了八個評估套件,發現當前的基準測試對於高頻危害類別是足夠的,但對於罕見且災難性的類別,則短少數個數量級。安全性基準測試並非不具資訊性。它們的資訊性是關於一組特定且可計算的命題;而它們所需要的紀律,便是明確指出是哪些命題。
English
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.