智能體如何在自動研究中失敗:對100項真實世界前沿研究任務的端到端診斷性評估
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
August 14, 2026
作者: Yanlin Fei, Nazhou Liu, Xinmiao Yu, Shaolong Chen, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das
cs.AI
摘要
人工智慧長久以來協助科學研究,但大型語言模型與代理式支架的快速進展正重塑此一格局;單一系統現已能執行從初始假設到最終發表論文的全階段研究,此範式現被稱為「自動研究」(AutoResearch)。現有評估對這些智能體如何運作或在哪裡失敗所揭露的資訊甚少。任務範圍狹窄,評估測量性能而非過程,而失敗診斷缺乏系統性的涵蓋範圍或產物層級的可視性。為填補此缺口,我們提出 AutoResearchEval,其包含 100 項以已發表的前沿科學為基礎的任務,涵蓋 7 個科學領域與完整的研究生命週期,包括構思、檢索、執行、分析、寫作與審查。評估 8 種支架—模型組合,產生 800 條自動研究智能體的軌跡,並附有過程層級的註釋。我們將這些洞察組織成「自動研究失敗分類法」(AutoResearch Failure Taxonomy, ARFT),這是一個包含 45 種以實證為基礎的失敗模式的框架。為實現可擴展的細粒度歸因,我們利用一個經人類校準的「智能體即評審」流程來檢查完整軌跡與中間產物。失敗模式匯聚到一個單一的總體限制:目前的智能體缺乏後設認知迴圈,亦即缺乏將自身產出與所發現的事實進行比對、在對不上時進行修訂,以及質疑自身所採取路徑是否穩健的能力。相同的模式在全部 8 種支架—模型組合中重複出現,包括受測的最強模型;這將缺陷定位於模型層級,而非任何特定的支架。編排層級的介入能否彌補此缺陷,是一個本工作未測試的開放問題。我們公開釋出 AutoResearchEval 與 ARFT,以促進自主科學發現領域的持續研究與發展。
English
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.