EVOHARNESSBENCH:你的代理能否跟上不斷演進的評估框架?
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
September 3, 2026
作者: Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
cs.AI
摘要
現代基於 LLM 的代理透過由工具、可重用技能和專業代理組成的框架運作,該框架塑造了它們所觀察到的內容以及它們能做的事。在實務中,隨著新能力的加入,這個框架會不斷演變。我們提出 EVOHARNESSBENCH,一個用於在受控框架演化下評估代理的基準測試,涵蓋三個軸線(工具、技能和代理)。與現有的代理持續學習基準測試不同,後者通常將非平穩性(即隨時間變化的內容)置於任務流中,同時保持框架固定不變,而 EVOHARNESSBENCH 將非平穩性置於外部提供的框架本身。它包含 17 個多階段框架流,由基於驗證器的基準測試確定性地構建而成,包含 802 個任務、520 個工具、42 項技能和 62 個代理。我們評估了兩種互補的設定,對應於框架演化的核心挑戰:部署評估,其隔離出隨著框架擴展時先前可存取能力的保留情況;以及自我演化適應評估,其測試累積經驗在新能力引入時是否仍然有用。我們的結果揭示了三個持續存在的差距。首先,僅框架擴展本身就可能降低先前已解決任務的效能,產生框架引發的遺忘。其次,自我演化適應所帶來的收益在框架演化的不同階段、能力軸線和環境中仍然不一致。第三,保留與適應可能朝不同方向拉扯:保留早期能力不一定能改善對新引入能力的適應,反之亦然。這些結果確立了框架演化為一項獨特的挑戰,對於構建能夠跟上不斷演變的框架、同時保留先前有效行為的代理而言。
English
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.