점수만으로는 발견을 입증할 수 없다: AI 연구 에이전트 감사를 위한 발견 인증 프로토콜
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
September 7, 2026
저자: Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng
cs.AI
초록
AI 연구 에이전트는 사전 지식, 공개 출처, 실험 피드백을 결합하여 유용한 결과를 산출한다. 발견 인증 프로토콜(DCP)은 이러한 결과에 관한 주장을 실행 가능한 복원 및 피드백 테스트로 전환한다. 게이트 1은 봉인된 평가에서 유용한 개선을 검증한다. 게이트 2는 매칭된 에이전트에게 등록된 시작 정보와 관찰된 웹 콘텐츠를 제공하는 한편, 목표 연구 이력은 제공하지 않는다. 수치 목표에 도달한 모든 유효한 방법은 복원 증인(recovery witness)을 제공하고 Core 거부권을 발동한다. DCP Core는 적절한 통제, 관찰된 복원 0건, 그리고 하나의 새로 등록된 에피소드에서 복원에 대한 유한 표본 경계를 요구한다. 선택적 게이트 3은 공유 체크포인트로부터 지정된 중립 정책에 상대적인 진실한 피드백의 평균 효과를 측정한다. DCP Evidence는 독립적 널 보정과 등록된 효과 마진 이후에 이 효과를 추가한다. 두 개의 통제된 감사는 서로 다른 모델에서 SQLite 최적화와 가상 촉매 제어에서 전체 프로토콜을 실행한다. 각각은 96개 에피소드에서 복원 0건을 산출했으며, 상한은 0.0468이었다. 각 짝지은 연구는 60쌍 널 연구를 통과하면서 30건의 진실한 복원과 0건의 중립 복원을 산출했다. 추가 사례들은 Core, 복원됨(recovered), 감사 미완료(audit-incomplete) 결정을 시험한다. 결정론적이며 LLM을 사용하지 않는 검증기는 동결된 증거로부터 결정을 재현한다. DCP는 AI 연구 전반에 걸쳐 유용한 결과, 대안 경로, 피드백 효과를 위한 공통 증거 언어를 제공한다.
English
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.