최종 점수를 넘어: 장기적 AI 연구 및 개발을 위한 에이전트의 체계적 평가
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
August 13, 2026
저자: Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
cs.AI
초록
자율 에이전트는 장기적인 실험을 통해 모델, 시스템 및 기타 기술 산출물을 개선하는 능력이 점점 더 향상되고 있다. 그러나 이러한 능력의 현재 상태를 이해하려면 평가가 최종 점수를 넘어서야 한다. 최종 점수는 진전이 어디에서 얻어지거나 상실되는지 드러내지 않으며, 축적된 경험이 이후의 결정을 개선하는지도 나타내지 않기 때문이다. 따라서 우리는 솔루션 프레이밍, 실행, 피드백 제어를 통해 실행 내 행동을 특성화하는 규칙 기반 지표와 작업 내 및 작업 간 경험 재사용을 평가하는 통제된 비교를 활용하는 새로운 프레임워크를 기반으로 36개의 장기 과제에 걸쳐 7개의 최첨단 모델에 대한 체계적인 평가를 제시한다. 결과는 현재 에이전트가 완전히 자율적인 연구자라기보다는 공학적 최적화 도구에 가깝게 작동함을 보여준다. 즉, 실용적인 솔루션을 공식화하고 구현할 수 있지만, 성능은 실행마다 크게 달라지며, 가장 강력한 솔루션은 주로 기존 기법을 변형하거나 결합한 것이고, 진정한 방법론적 혁신은 여전히 드물다. 상세 분석은 관찰된 성능이 여러 요인에 의해 형성됨을 보여준다. 유사한 최종 결과 이면의 서로 다른 프로세스 병목, 후속 결정을 돕거나 오도할 수 있는 경험 재사용, 그리고 성능 안정성에 영향을 미치는 하네스 설계가 그러한 요인에 포함된다. 이러한 발견은 모델 훈련, 추론 시간 전략, 경험 관리 및 하네스 설계를 개선하기 위한 구체적인 방향을 제시한다.
English
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.