ChatPaper.aiChatPaper

AI 레드팀 평가가 증명할 수 있는 것과 증명할 수 없는 것.

What AI Red-Team Evaluations Can and Cannot Prove

July 23, 2026
저자: Bandana Kaur
cs.AI

초록

레드팀 평가(red-team evaluation)는 일부 주장을 지지하지만 다른 주장은 지지하지 않으며, 이 둘의 경계는 단순한 판단의 문제가 아니라 계산 가능하다. 우리는 평가의 증거 상한(evidential ceiling)을 고정된 테스트 예산 하에서 하나의 결과가 믿음(belief)을 이동시킬 수 있는 최대 배수로 정의하고, 벤치마크 귀무 결과(null result)에 대해 닫힌 형태(closed form)로 유도한 다음, 이를 사용하여 그 경계를 정확히 위치시킨다. 우리는 계산 가능한 피해율(harm rate) 이상에서는 적절한 규모의 벤치마크가 명시된 증거 기준(evidentiary standard)에 따라 해당 범주를 인증하며, 이때 완전 무결과(clean sheet)는 두 가지 가능한 관측 중 더 강력한 증거가 되어 단일하게 재현된 실패보다 우월함을 발견한다. 그 피해율 이하에서는, 고정된 채점 규칙과 대략적으로 독립적인 시행 구조 하에서, 실현 가능한 규모의 어떠한 수동적 벤치마크도 명시된 안전성 증거를 제공하지 못한다. 두 체제 사이의 교차점은 닫힌 형태를 갖는다. 이 경계는 벤치마크에 국한되지 않는다. 절차의 가설 조건부 도출률(hypothesis conditioned elicitation rate)로 표현될 때, 이는 적응형 및 자동화된 레드팀링도 포괄하며, 공격 성공이 아니라 가설 간의 변별(discrimination)이 증거적 가치를 결정함을 보여준다. 여덟 개의 평가 스위트(suite)를 이 경계에 대해 감사한 결과, 현재의 벤치마크는 고빈도 피해 범주에는 적합하지만, 희귀하고 치명적인 범주에 대해서는 수 자릿수만큼 부족함을 발견한다. 안전성 벤치마크는 정보를 제공하지 않는 것이 아니다. 그것들은 특정하고 계산 가능한 명제 집합에 대해 정보를 제공하며, 요구되는 규율은 어떤 명제에 대해 정보를 제공하는지를 명시하는 것이다.
English
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.