ChatPaper.aiChatPaper

Evo-Bench:語言模型能否改進智能體框架?

Evo-Bench: Can Language Models Improve Agent Harness?

August 10, 2026
作者: Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang
cs.AI

摘要

大型語言模型(LLMs)已推動自主代理(autonomous agents)的快速進展,然而標準評估仍局限於靜態任務求解。一個新興的前沿領域是框架演化(harness evolution)——即代理自主優化其自身操作框架的能力。然而,系統性地基準測試這項能力仍具挑戰性,因為現有評估要麼無法將框架改進與基礎模型強度分離,要麼無法防止針對特定任務的過擬合,要麼無法捕捉長時程的迭代研究。為應對這些挑戰,我們提出了 Evo-Bench,這是首個專門設計來評估模型在搜尋(Search)、辦公室(Office)與通用代理(General agent)領域中內在框架演化能力的基準。為嚴格隔離此能力,Evo-Bench 採用了一種新穎的框架引導式建構框架:它利用輔助任務演化來識別對框架改進真正敏感的任務,接著透過敏感度感知的分層拆分(sensitivity-aware stratified splitting)來確保跨套件的穩健泛化。對九個前沿與開放權重模型的廣泛評估顯示,頂尖模型實現了高達 16.6 分的巨幅絕對增益,幾乎逼近最先進的人工設計基準線。至關重要的是,雖然自主演化在通用任務中優於人工框架,並在搜尋任務中表現出色,但在要求高度特定處理流程的辦公室任務中則表現吃力。此外,我們的分析揭露了諸如早期飽和等關鍵時間異常,同時證明所合成的框架可作為高度可遷移的推理結構,持續提升多種策略模型(policy models)的表現。
English
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.