AIエージェントは非定型のAI研究を実施できるか?——2つのケーススタディからの初期の証拠
Can AI agents conduct open-ended AI research? Early evidence from two case studies
July 29, 2026
著者: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
cs.AI
要旨
爆発的なAI進歩の予測は、AIエージェントがAI研究を自動化することに依存している。しかし、エージェントが開かれたAI研究を遂行できるかどうかについての証拠は乏しい。現在の評価手法は、検証可能な狭いタスクでエージェントをテストするもの(開かれた研究を除外する)か、AIが生成した論文を盲検査読に提出するもの(過負荷で確率的、かつ査読品質が低い)に限られる。我々は、AI研究開発の自動化に向けた進歩を測定する第三の方法を導入する。エージェントは、未発表の質の高い論文の中心的かつ開かれた研究課題に取り組み、その論文の原著者がエージェントの出力を評価する。我々はこれをシャドウ評価と呼ぶ。未発表のNeurIPS 2026投稿論文2本に対してシャドウ評価を実施し、最先端エージェントに6日間と数千ドルの計算リソースを与えた。エージェントはすべてのエンジニアリング作業を人間の助けなしで完了したが、研究課題に回答するための実質的な進展は達成できなかった。その結果、両論文とも原著者により明白に却下された。我々は5つの再発する失敗パターンを特定した:出版可能な研究の基準に関する判断力の低さ、研究デザインの欠点に対する創造性のない対応、行き詰まりからの非効率的な後退、リソース認識の低さ、指示の逸脱である。別のモデルとスキャフォールドを用いたロバスト性チェックでも、これらの失敗が再現された。我々は専門家による査読結果、調査回答、エージェントのリポジトリ、ログを公開する。本研究の結果は、今日のエージェントがAI研究のエンジニアリングを実行できる一方、研究ライフサイクルの重要な部分に苦慮しているという初期の証拠を提供する。
English
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.