僅憑分數不足以證明發現:稽核 AI 研究代理的發現認證協議
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
September 7, 2026
作者: Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng
cs.AI
摘要
人工智慧研究代理結合先驗知識、公開來源與實驗回饋,以產生有用的結果。發現認證協議(DCP)將關於這些結果的主張轉化為可執行的復原與回饋測試。Gate 1 驗證在密封評估上的有用改善。Gate 2 給予配對代理註冊的起始資訊與觀察到的網路內容,同時扣留目標研究歷史。每個達到數值目標的有效方法都會提供復原見證,並觸發 Core 否決。DCP Core 要求充分的控制、零觀察到的復原,以及在一個全新註冊回合中復原的有限樣本界限。可選的 Gate 3 衡量真實回饋相對於來自共享檢查點的指定中性政策的平均效應。DCP Evidence 在獨立虛無校準與註冊效應邊際之後加入此效應。兩個受控稽核在不同模型下,於 SQLite 最佳化與虛擬催化劑控制中演練完整協議。每個在 96 個回合中產生零次復原,上限為 0.0468。每項配對研究產生 30 次真實復原與零次中性復原,並通過 60 對虛無研究。額外案例演練 Core、已復原與稽核不完整決策。一個確定性、不依賴 LLM 的驗證器從凍結證據重現這些決策。DCP 為 AI 研究中的有用成果、替代路徑與回饋效應提供共同的證據語言。
English
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.