ChatPaper.aiChatPaper

AgentCompass:智能體能力的統一評估基礎設施

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

July 15, 2026
作者: Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
cs.AI

摘要

隨著大型語言模型(LLMs)發展為自主代理,建立統一的評估基礎設施變得至關重要。然而,當前的評估流程仍然高度分散且緊密耦合,阻礙了可重複性並導致冗餘的工程開發。為解決此問題,我們提出AgentCompass——一個開源、輕量且可擴展的評估基礎設施,專為基於LLM的代理設計。AgentCompass將評估流程組織為三個獨立組件:基準(Benchmark)、測試框架(Harness)與環境(Environment),從而實現靈活的配置,無需重新實現複雜的執行邏輯。此外,它具備容錯的非同步運行環境以及全面的軌跡分析工具,能透明地診斷如獎勵欺騙(reward-hacking)等細微的失敗模式。原生支援橫跨五個能力維度的20多個基準測試,AgentCompass為社群提供了一個可擴展且可重複的基礎設施,以推動代理研究的進展。
English
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.