ChatPaper.aiChatPaper

最終スコアを超えて:長期的AI研究開発のためのエージェントの系統的評価

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

August 13, 2026
著者: Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
cs.AI

要旨

自律エージェントは、モデル、システム、その他の技術的成果物を、長期的な実験を通じて改善する能力を急速に高めつつある。しかし、この能力の現状を理解するには、評価が最終スコアだけに留まってはならない。最終スコアは、進歩がどこで得られ、どこで失われているのかを明らかにせず、また、蓄積された経験がその後の意思決定を改善するかどうかも示さないからである。そこで本稿では、実行中の挙動をルールベースの指標によって特徴づける新しいフレームワークに基づき、7つのフロンティアモデルを36の長期的タスクで系統的に評価する。このフレームワークは、解法の枠組み(Solution Framing)、実行(Execution)、フィードバック制御(Feedback Control)を通じて実行中の挙動を特徴づけ、統制比較によってタスク内およびタスク間の経験再利用を評価する。結果は、現在のエージェントが完全に自律的な研究者というよりも、むしろ工学的なオプティマイザーとして機能していることを示している。すなわち、実用的な解法を定式化し実装することはできるが、その性能は実行ごとに大きくばらつき、最も優れた解法は主に既存の技法を適応または組み合わせたものであり、真に方法論的な新規性は依然として稀である。詳細な分析により、観測される性能は複数の要因によって形成されることが明らかになった。類似した最終結果の背後にある異なるプロセス上のボトルネック、その後の意思決定を助けたり誤らせたりする経験再利用、そして性能の安定性に影響を与えるハーネス設計などである。これらの知見は、モデル訓練、推論時戦略、経験管理、およびハーネス設計を改善するための具体的な方向性を示唆している。
English
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.