ChatPaper.aiChatPaper

Evo-Bench: 言語モデルはエージェントハーネスを改善できるか?

Evo-Bench: Can Language Models Improve Agent Harness?

August 10, 2026
著者: Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang
cs.AI

要旨

大規模言語モデル(LLM)は自律エージェントの急速な進歩を牽引してきたが、標準的な評価手法は依然として静的なタスク解決に限定されている。新たな研究フロンティアとして浮上しているのが、ハーネス進化——エージェントが自身の動作ハーネスを自律的に最適化する能力——である。しかし、この能力を体系的にベンチマークすることは依然として困難であり、既存の評価手法は、ハーネスの改善をベースモデルの性能から分離できず、タスク固有の過学習を防止できず、長期的な反復研究を捉えることもできない。これらの課題に対処するため、我々はSearch、Office、Generalエージェント領域におけるモデルの本質的なハーネス進化能力を評価する初のベンチマークであるEvo-Benchを提案する。この能力を厳密に分離するため、Evo-Benchは新規のハーネス誘導型構築フレームワークを採用する。具体的には、補助タスク進化を活用してフレームワークの改善に真に敏感なタスクを特定し、続いて感度認識型層化分割を適用して頑健なクロススイート汎化を保証する。9つのフロンティアモデルおよびオープンウェイトモデルにわたる広範な評価から、トップモデルは最大16.6ポイントの絶対的改善を達成し、最先端の人間設計ベースラインに肉薄することが明らかになった。重要なことに、自律的進化はGeneralタスクにおいて人工ハーネスを上回り、Searchタスクでも優れた性能を発揮する一方、高度に特化した処理フローを要求するOfficeタスクでは苦戦する。さらに、我々の分析は早期飽和などの重大な時間的異常を浮き彫りにするとともに、合成されたハーネスが転移可能性の高い推論構造として機能し、多様なポリシーモデルを一貫して向上させることを実証している。
English
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.