AI 에이전트가 개방형 AI 연구를 수행할 수 있는가? 두 사례 연구를 통한 초기 증거
Can AI agents conduct open-ended AI research? Early evidence from two case studies
July 29, 2026
저자: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
cs.AI
초록
폭발적인 AI 발전에 대한 예측은 AI 에이전트가 AI 연구를 자동화하는 것에 달려 있다. 그러나 에이전트가 개방형 AI 연구를 수행할 수 있는지에 대한 증거는 부족하다. 현재 평가 방식은 좁고 검증 가능한 작업에 한정된 에이전트를 테스트하여 개방형 연구를 배제하거나, AI가 작성한 논문을 블라인드 피어 리뷰에 제출하는 방식을 사용하는데, 이는 과도하게 확장되어 있고 결과가 불확실하며 리뷰 품질이 낮은 문제를 안고 있다. 우리는 AI R&D 자동화를 향한 진전을 측정하는 세 번째 방법을 제안한다. 에이전트는 고품질의 미발표 논문의 핵심적이고 개방적인 연구 질문을 맡고, 논문의 원저자들이 그 결과물을 평가한다. 우리는 이를 섀도 평가(shadow evaluations)라고 부른다. 우리는 미발표된 NeurIPS 2026 제출 논문 두 편에 대해 섀도 평가를 실시했으며, 최첨단 에이전트에게 6일의 시간과 수천 달러 상당의 컴퓨팅 자원을 제공했다. 에이전트는 인간의 도움 없이 모든 엔지니어링 작업을 완료했지만, 연구 질문에 대한 실질적인 진전을 이루지 못했다. 그 결과, 두 논문 모두 원저자들에 의해 명백히 거절되었다. 우리는 다섯 가지 반복적인 실패 패턴을 식별했다: 게재 가능한 연구 수준에 대한 판단 부족, 연구 설계의 한계에 대한 비창의적 대응, 막다른 골목에서의 비효과적 회귀, 자원 인식 부족, 지시 방황(instruction drift)이다. 두 번째 모델과 스캐폴드를 사용한 강건성 검증에서도 이러한 실패가 재현되었다. 우리는 전문가 리뷰, 설문 응답, 에이전트 저장소 및 로그를 공개한다. 우리의 결과는 오늘날의 에이전트가 AI 연구의 엔지니어링을 수행할 수 있지만, 연구 생애 주기의 핵심 부분에서는 어려움을 겪는다는 초기 증거를 제공한다.
English
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.