ChatPaper.aiChatPaper

AI智能体能否进行开放式AI研究?来自两个案例研究的初步证据

Can AI agents conduct open-ended AI research? Early evidence from two case studies

July 29, 2026
作者: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
cs.AI

摘要

對人工智慧爆炸性進展的預測,往往建立在AI Agent能自動化進行AI研究的基礎上。然而,關於AI Agent能否執行開放式AI研究的證據仍然薄弱。現有的評估方式,要麼將Agent侷限於狹隘、可驗證的任務上,從而排除了開放式研究;要麼將AI生成的論文提交至雙盲同儕審查,而這種審查機制負擔過重、充滿隨機性,且審稿品質不佳。我們提出第三種衡量AI研發自動化進展的方法:讓AI Agent接手一篇高品質未發表論文的中心開放式研究問題,並由該論文的原始作者對其輸出進行評分。我們將此方法稱為影子評估。我們針對兩篇未發表的NeurIPS 2026投稿論文進行了影子評估,賦予前沿Agent六天的時間與數千美元的計算資源。這些Agent在沒有人類協助的情況下完成了所有工程任務,但在解答研究問題的核心方向上卻未能取得實質進展。因此,這兩篇論文均被原始作者明確拒絕。我們歸納出五種反覆出現的失敗模式:對於可發表研究的門檻判斷不佳、對研究設計缺陷缺乏創意回應、無法有效從死胡同中回頭、資源意識薄弱,以及指令漂移。我們使用第二組模型與框架進行穩健性檢驗,結果同樣再現了這些失敗。我們公開了專家評審意見、問卷調查結果、Agent程式碼儲存庫以及完整日誌。我們的研究結果提供了早期證據,證明當今的AI Agent能夠勝任AI研究的工程部分,但在研究生命週期的關鍵環節上仍有諸多困難。
English
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.