ChatPaper.aiChatPaper

超越最终得分:面向长时程AI研究与开发的智能体系统性评估

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

August 13, 2026
作者: Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
cs.AI

摘要

自主智能体正日益具备通过长时域实验来改进模型、系统及其他技术制品的能力。然而,要理解这一能力的当前状态,评估必须超越最终得分——最终得分既无法揭示进展得失发生在何处,也无法表明累积经验是否改善后续决策。为此,我们基于一个新框架,对7个前沿模型在36项长时域任务上进行了系统性评估。该框架采用基于规则的评价指标,通过问题构架、执行和反馈控制三个维度刻画运行内行为,并通过受控比较评估任务内及任务间的经验复用。结果表明,当前智能体的行为更接近工程优化器而非完全自主的研究者:它们能够提出并实施实用解决方案,但其表现在不同运行之间差异显著;它们最强的解决方案主要是在改编或组合已有技术,而真正的方法论创新仍然罕见。详细分析揭示,观察到的表现受多种因素共同塑造,包括相似最终结果背后不同的过程瓶颈、既可能帮助也可能误导后续决策的经验复用,以及影响表现稳定性的框架设计。这些发现为改进模型训练、推理时策略、经验管理和框架设计提供了具体方向。
English
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.