エージェントはAutoResearchでどのように失敗するのか:100件の実世界の最先端研究タスクにおけるエンドツーエンド診断的評価
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
August 14, 2026
著者: Yanlin Fei, Nazhou Liu, Xinmiao Yu, Shaolong Chen, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das
cs.AI
要旨
AIは長きにわたり科学研究を支援してきたが、大規模言語モデル(LLM)とエージェント的スキャフォールドの急速な進歩により、その状況は一変しつつある。現在では、単一のシステムが初期仮説から最終的な論文発表に至るまで研究の全段階を遂行できるようになっており、このパラダイムはAutoResearchと呼ばれている。既存の評価では、これらのエージェントがどのように動作し、どこで破綻するのかについてほとんど明らかにされていない。タスクの範囲は狭く、評価はプロセスではなく性能を測定しており、失敗の診断には体系的な網羅性や成果物レベルの可視性が欠けている。このギャップに対処するため、我々はAutoResearchEvalを導入する。これは、7つの科学領域にわたり発表された最先端の科学に基づく100のタスクと、発想、検索、実行、分析、執筆、レビューを含む研究ライフサイクル全体を特徴とする。8つのハーネス・モデル組み合わせの評価により、プロセスレベルのアノテーションを備えた800のAutoResearchエージェント軌跡が得られた。我々はこれらの洞察を、45の経験的に基づく失敗パターンの枠組みであるAutoResearch Failure Taxonomy(ARFT)として体系化した。スケーラブルで詳細な帰属を可能にするため、我々は人間により較正されたエージェント・アズ・ジャッジのパイプラインを活用し、完全な軌跡と中間成果物を検査する。失敗パターンは、単一の包括的な限界に収束する。すなわち、現在のエージェントにはメタ認知ループが欠如しているということである。メタ認知ループとは、得られた結果と照らし合わせて自らが生成したものを検証し、それが妥当でない場合に修正し、取った経路が適切であったかを問い直す能力を伴う。同じパターンは、テストされた最も強力なモデルを含むすべての8つのハーネス・モデル組み合わせにわたって再発し、その欠陥は特定のスキャフォールドではなくモデルレベルに位置づけられる。オーケストレーションレベルの介入がこの欠陥を解消できるかどうかは、本研究では検証されていない未解決の問いである。我々は、自律的な科学的発見における研究開発の継続を促進するため、AutoResearchEvalとARFTを公開する。
English
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.