AgentCompass: エージェント能力のための統一評価基盤
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
July 15, 2026
著者: Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
cs.AI
要旨
大規模言語モデル(LLM)が自律エージェントへと進化するにつれ、統一的な評価基盤の重要性が高まっている。しかし現状の評価パイプラインは依然として非常に断片的かつ密結合であり、再現性を妨げるとともに冗長なエンジニアリングを引き起こしている。この課題に対処するため、我々はLLMベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能な基盤であるAgentCompassを提案する。AgentCompassは評価プロセスをBenchmark、Harness、Environmentという3つの独立したコンポーネントに整理することで、複雑な実行ロジックの再実装を必要とせずに柔軟な設定を可能にする。さらに、耐障害性のある非同期ランタイムと包括的な軌跡分析ツールを備え、報酬ハッキングのような微妙な障害モードを透過的に診断できる。AgentCompassは5つの能力次元にわたる20以上のベンチマークをネイティブでサポートし、エージェント研究を前進させるためのスケーラブルで再現可能な基盤をコミュニティに提供する。
English
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.