超越最終得分:針對長時程人工智慧研究與開發之智慧體系統性評測
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
August 13, 2026
作者: Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
cs.AI
摘要
自主代理日益具備透過長期實驗來改進模型、系統及其他技術產品的能力。然而,欲掌握此能力之現狀,評估不應僅止於最終分數,因為最終分數既無法揭示進展在何處獲得或喪失,也無法反映累積的經驗是否有助於後續決策。因此,我們基於一個新框架,對七個前沿模型在三十六項長期任務上進行系統性評估。該框架採用基於規則的指標,透過「解決方案構思」、「執行」與「回饋控制」三個面向來刻畫運行內行為,並利用受控比較來評估任務內與任務間的經驗重用。結果顯示,當前代理的運作模式更接近工程優化器,而非全自主研究人員:它們能構思並實作實用的解決方案,但其表現在不同運行間存在大幅差異;其最強解決方案主要是在改編或結合既有技術;真正的方法論創新仍屬罕見。詳細分析揭示,觀察到的表現受多重因素影響,包括相似最終結果背後截然不同的流程瓶頸、可能助益亦可能誤導後續決策的經驗重用,以及影響表現穩定性的框架設計。這些發現為改進模型訓練、推論時期策略、經驗管理與框架設計提供了具體方向。
English
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.