EVOHARNESSBENCH:あなたのエージェントは進化するハーネスに追随できるか?
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
September 3, 2026
著者: Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
cs.AI
要旨
現代のLLMベースのエージェントは、ツール、再利用可能なスキル、専門エージェントからなるハーネスを通じて動作し、そのハーネスがエージェントの観測対象と実行可能な行動を形作る。実際には、このハーネスは新たな能力が追加されるにつれて継続的に進化する。我々は、3つの軸(ツール、スキル、エージェント)にわたる制御されたハーネス進化の下でエージェントを評価するためのベンチマーク、EVOHARNESSBENCHを導入する。エージェント向けの既存の継続学習ベンチマークは、通常、ハーネスを固定したままタスクストリームに非定常性(すなわち、時間とともに変化するもの)を置くのに対し、EVOHARNESSBENCHは、外部から与えられるハーネスそのものに非定常性を置く。これは、検証器ベースのベンチマークから決定論的に構築された17の多段階ハーネスストリームを含み、802タスク、520ツール、42スキル、62エージェントから構成される。我々は、ハーネス進化の中心的課題に対応する2つの相補的設定を評価する。すなわち、ハーネスが拡大するにつれて以前に利用可能だった能力の保持を分離するデプロイ評価と、新たな能力が導入される中で蓄積された経験が依然として有用かを検証する自己進化型適応評価である。我々の結果は3つの根強いギャップを明らかにする。第一に、ハーネスの拡大それ自体が、以前に解決したタスクの性能を低下させ得るため、ハーネス誘発性忘却を生じさせる。第二に、自己進化型適応による利得は、ハーネス進化の段階、能力軸、環境をまたいで一貫しない。第三に、保持と適応は異なる方向に作用し得る。すなわち、以前の能力を保持することは、新たに導入された能力への適応を必ずしも改善せず、その逆も同様である。これらの結果は、進化するハーネスに追随しつつ以前に有効だった行動を保持できるエージェントを構築する上で、ハーネス進化が独自の課題であることを確立する。
English
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.