人工智能智能体能开展开放式AI研究吗?基于两项案例研究的早期证据
Can AI agents conduct open-ended AI research? Early evidence from two case studies
July 29, 2026
作者: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
cs.AI
摘要
对爆炸性AI进展的预测,依赖于AI研究员自动化的智能体。但关于智能体能否执行开放性AI研究的证据仍然薄弱。当前评估方法要么局限于狭窄的可验证任务(此类任务排除开放性研究),要么将AI生成的论文提交给盲审同行评审——这种评审体系已不堪重负、具有随机性且评审质量低下。我们提出第三种衡量AI研发自动化进展的路径:让智能体接手某篇高质量未发表论文的核心开放性研究问题,并由该论文的原作者对其输出进行评分。我们称之为影子评估。我们针对两篇未发表的NeurIPS 2026投稿论文开展了影子评估,为前沿智能体提供六天时间及数千美元的计算资源。智能体在无人协助下完成了所有工程任务,却未能在解答研究问题方面取得实质性进展。结果,两篇论文均被作者明确拒绝。我们识别出五个反复出现的失败模式:对可发表研究成果标准的判断力薄弱、对研究设计缺陷的应对缺乏创造性、从死胡同中有效回退的能力不足、资源意识薄弱以及指令漂移。使用第二种模型与脚手架进行的稳健性检验再现了这些失败。我们公开了专家评审意见、调查问卷反馈、智能体代码仓库及运行日志。研究结果表明:当前智能体能够完成AI研究的工程部分,但在研究生命周期的关键环节仍存在重大短板。
English
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.