ChatPaper.aiChatPaper

Evo-Bench: 언어 모델이 에이전트 하네스를 개선할 수 있는가?

Evo-Bench: Can Language Models Improve Agent Harness?

August 10, 2026
저자: Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang
cs.AI

초록

대규모 언어 모델(LLM)은 자율 에이전트 분야의 급속한 발전을 견인해 왔으나, 표준 평가 방식은 여전히 정적 과제 해결에 국한되어 있다. 새로운 연구의 최전선은 하네스 진화(harness evolution), 즉 에이전트가 자신의 운영 하네스를 자율적으로 최적화하는 능력이다. 그러나 기존 평가 방식은 하네스 개선을 기본 모델의 성능과 분리하지 못하거나, 과제별 과적합을 방지하지 못하거나, 장기적 반복 연구를 포착하지 못하는 한계가 있어 이 능력을 체계적으로 벤치마킹하는 것은 여전히 어려운 과제로 남아 있다. 이러한 문제를 해결하기 위해, 우리는 검색(Search), 오피스(Office), 일반(General) 에이전트 도메인에서 모델의 본질적인 하네스 진화 능력을 평가하도록 설계된 최초의 벤치마크인 Evo-Bench를 소개한다. 이 능력을 엄밀하게 분리 평가하기 위해, Evo-Bench는 새로운 하네스 기반 구축 프레임워크를 채택한다. 이 프레임워크는 보조 과제 진화(auxiliary-task evolution)를 활용하여 프레임워크 개선에 진정으로 민감한 과제를 식별하고, 이어서 민감도 인식 계층 분할(sensitivity-aware stratified splitting)을 적용하여 강건한 교차 스위트 일반화를 보장한다. 아홉 개의 최첨단 및 오픈 가중치 모델에 대한 광범위한 평가 결과, 최상위 모델은 16.6포인트에 달하는 큰 절대적 성능 향상을 달성하여 최첨단 수작업 설계 기준선에 근접하는 것으로 나타났다. 중요한 점은, 자율 진화가 일반 과제에서 수작업 하네스를 능가하고 검색 과제에서 탁월한 성과를 보이는 반면, 고도로 특화된 처리 워크플로우를 요구하는 오피스 과제에서는 어려움을 겪는다는 것이다. 또한 우리의 분석은 조기 포화와 같은 중대한 시간적 이상 현상을 드러내는 동시에, 합성된 하네스가 전이 가능성이 높은 추론 구조로 작용하여 다양한 정책 모델을 일관되게 향상시킴을 입증한다.
English
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.