ChatPaper.aiChatPaper

AutoResearch에서 에이전트는 어떻게 실패하는가: 100가지 실제 최첨단 연구 과제에 대한 종단 간 진단 평가

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

August 14, 2026
저자: Yanlin Fei, Nazhou Liu, Xinmiao Yu, Shaolong Chen, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das
cs.AI

초록

AI는 오랫동안 과학 연구를 지원해 왔지만, 대규모 언어 모델(LLM)과 에이전트형 스캐폴드(agentic scaffolds)의 급속한 발전이 연구 환경을 재편하고 있다. 이제 단일 시스템이 초기 가설 수립부터 최종 논문 게재까지 연구의 전 과정을 수행할 수 있으며, 이러한 패러다임은 AutoResearch라고 불린다. 기존 평가는 이러한 에이전트가 어떻게 작동하는지, 어디에서 실패하는지에 대해 거의 밝혀내지 못한다. 과제의 범위가 좁고, 평가는 과정이 아닌 성과만 측정하며, 실패 진단은 체계적 포괄성이나 산출물 수준의 가시성이 부족하다. 이러한 격차를 해소하기 위해 우리는 AutoResearchEval을 제안한다. AutoResearchEval은 7개 과학 분야의 출판된 최첨단 연구에 기반한 100개 과제로 구성되며, 아이디어 구상, 검색, 실행, 분석, 작성, 검토를 포함한 연구 수명주기 전반을 다룬다. 8가지 하네스-모델 조합을 평가하여 프로세스 수준 주석이 포함된 800개의 AutoResearch 에이전트 궤적을 확보했다. 우리는 이러한 통찰을 45개의 경험적으로 뒷받침된 실패 패턴으로 구성된 프레임워크인 AutoResearch 실패 분류 체계(ARFT)로 체계화했다. 확장 가능한 세밀한 원인 규명을 위해, 인간 보정된 에이전트-판사(agent-as-a-judge) 파이프라인을 활용하여 전체 궤적과 중간 산출물을 검사한다. 실패 패턴은 단일한 포괄적 한계로 수렴한다. 즉, 현재 에이전트는 메타인지 루프가 부족하다는 것이다. 메타인지 루프란 자신이 발견한 내용과 자신이 생성한 결과물을 대조하고, 그것이 타당하지 않을 때 수정하며, 자신이 취한 경로가 건전했는지 의문을 제기하는 능력을 의미한다. 동일한 패턴은 테스트된 가장 강력한 모델을 포함한 8개 하네스-모델 조합 모두에서 반복되며, 이는 결함이 특정 스캐폴드에 있는 것이 아니라 모델 수준에 있음을 보여 준다. 오케스트레이션 수준의 개입이 이를 해결할 수 있는지는 본 연구에서 검증하지 않은 미해결 질문이다. 우리는 자율적 과학 발견에 대한 지속적인 연구와 개발을 촉진하기 위해 AutoResearchEval과 ARFT를 공개적으로 배포한다.
English
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.