AgentCompass: 에이전트 역량을 위한 통합 평가 인프라
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
July 15, 2026
저자: Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
cs.AI
초록
대규모 언어 모델(LLM)이 자율 에이전트로 진화함에 따라 통합된 평가 인프라의 필요성이 중요해지고 있습니다. 그러나 현재의 평가 파이프라인은 여전히 매우 파편화되어 있고 밀접하게 결합되어 있어 재현성을 저해하고 중복된 엔지니어링 작업을 유발합니다. 이러한 문제를 해결하기 위해, 우리는 LLM 기반 에이전트 평가를 위한 오픈소스, 경량화, 확장 가능한 인프라인 AgentCompass를 소개합니다. AgentCompass는 평가 과정을 Benchmark, Harness, Environment라는 세 가지 독립적인 구성 요소로 구성함으로써, 복잡한 실행 로직을 재구현하지 않고도 유연한 구성을 가능하게 합니다. 또한, 내결함성 비동기 런타임과 포괄적인 궤적 분석 도구를 제공하여 보상 해킹과 같은 미묘한 실패 모드를 투명하게 진단할 수 있습니다. 다섯 가지 능력 차원에 걸쳐 20개 이상의 벤치마크를 기본 지원하는 AgentCompass는 에이전트 연구 발전을 위한 확장 가능하고 재현 가능한 인프라를 커뮤니티에 제공합니다.
English
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.