EVOHARNESSBENCH: 당신의 에이전트는 진화하는 하네스와 보조를 맞출 수 있는가?
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
September 3, 2026
저자: Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty
cs.AI
초록
최신 LLM 기반 에이전트는 도구, 재사용 가능한 스킬, 전문 에이전트로 구성된 하네스를 통해 작동하며, 이 하네스는 에이전트가 무엇을 관찰하고 무엇을 할 수 있는지를 규정한다. 실제로 이러한 하네스는 새로운 능력이 추가됨에 따라 지속적으로 진화한다. 우리는 통제된 하네스 진화를 세 축(도구, 스킬, 에이전트)에 걸쳐 평가하기 위한 벤치마크인 EVOHARNESSBENCH를 소개한다. 일반적으로 하네스를 고정한 채 비정상성(즉, 시간에 따라 변하는 것)을 과제 스트림에 두는 기존의 에이전트 지속 학습 벤치마크와 달리, EVOHARNESSBENCH는 비정상성을 외부에서 제공되는 하네스 자체에 둔다. 이는 검증기 기반 벤치마크로부터 결정론적으로 구성된 17개의 다단계 하네스 스트림을 포함하며, 802개 과제, 520개 도구, 42개 스킬, 62개 에이전트로 이루어져 있다. 우리는 하네스 진화의 핵심 난제에 대응하는 두 가지 상호 보완적 설정을 평가한다: 배포 평가는 하네스가 확장됨에 따라 이전에 접근 가능했던 역량의 유지 여부를 분리해 측정하고, 자기 진화 적응 평가는 새로운 능력이 도입될 때 축적된 경험이 유용하게 남는지를 시험한다. 우리의 결과는 세 가지 지속적인 격차를 드러낸다. 첫째, 하네스 확장만으로도 이전에 해결한 과제에서 성능이 저하되어 하네스 유발 망각이 발생할 수 있다. 둘째, 자기 진화 적응에서 얻는 이점은 하네스 진화의 단계, 능력 축, 환경 전반에 걸쳐 일관되지 않다. 셋째, 유지와 적응은 서로 다른 방향으로 작용할 수 있다: 이전 역량을 보존하는 것이 새로 도입된 능력에 대한 적응을 반드시 향상시키지는 않으며, 그 반대도 마찬가지다. 이러한 결과는 진화하는 하네스를 따라가면서도 이전의 효과적인 행동을 보존할 수 있는 에이전트를 구축하는 데 있어 하네스 진화가 별개의 난제임을 확립한다.
English
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.