AI红队评估能证明什么与不能证明什么
What AI Red-Team Evaluations Can and Cannot Prove
July 23, 2026
作者: Bandana Kaur
cs.AI
摘要
对AI模型的红队评估支持某些主张而不支持另一些主张,而这二者之间的界限是可以计算的,而不仅仅是一个判断问题。我们将评估的证据上限定义为一个结果在固定测试预算下能使信念移动的最大倍数,并针对基准测试的零结果以闭式形式推导出该上限,从而精确地定位上述界限。我们发现,在可计算的危害率之上,一个中等规模的基准测试足以按既定证据标准对某一类别给出认证;此时,零失败记录是两种可能观察中更强的一种,其证据分量超过单个可复现的失败。在该危害率之下,任何可行规模的被动基准测试都无法在固定评分规则和近似独立的试验结构下提供所要求的安全性证据。两种状态之间的交叉点具有闭式形式。这一界限并非基准测试所特有:以程序的假设条件引出率来表述,它也适用于自适应和自动化红队测试,并表明决定证据价值的是假设之间的区分,而非攻击是否成功。我们依据这一界限审计了八个评估套件,发现当前基准测试对高频危害类别来说是充分的,但对罕见且灾难性的危害类别则短缺几个数量级。安全基准测试并非没有信息量;它们提供的是关于一组特定且可计算的命题的信息,而它们所需要的规范是明确说明这些命题是什么。
English
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.