SWE-Bench Pro Verified:軟體工程代理的可靠基準測試
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
September 8, 2026
作者: Pujun Zheng, Zixin Shang, Shufan Jiang, Wenhui Tian, Dongsheng Zhu, Zerun Ma, Dingbo Yuan, Qi Zhang
cs.AI
摘要
SWE-Bench Pro 已成為評估軟體工程代理在具挑戰性儲存庫層級任務上的標準基準測試。然而,我們的分析工作顯示,其評估受到兩種不可靠來源的損害:獎勵駭取行為,因標準解答或隱藏評估資訊外洩而得以發生;以及任務品質問題,包括具誤導性的問題陳述與範圍不當的測試。這些問題可能誇大基準測試表現,並掩蓋代理真正的程式設計能力。我們提出 SWE-Bench Pro Verified,這是 SWE-Bench Pro 的驗證版,可解決上述兩個問題。我們的方法結合了反駭取防護措施與任務精煉:前者可消除主要洩漏管道,而不干擾代理的正常功能;後者則以最小幅度修正有缺陷實例中的不一致之處。在 SWE-Bench Pro Verified 上的評估顯示,部分模型的表現明顯較先前報告的結果差,這意味著 SWE-Bench Pro 的既有結果可能高估真實的軟體工程能力。SWE-Bench Pro Verified 為評估軟體工程代理提供了更值得信賴的基準測試。
English
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.