超越成功率:成本感知的攻防安全代理评估
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
July 17, 2026
作者: Paul Kassianik, Blaine Nelson, Yaron Singer
cs.AI
摘要
安全代理评估通常会在充裕的推理预算下衡量其峰值攻击能力,重点关注漏洞发现、漏洞利用开发、渗透测试和CTF解题。这类评估虽有价值却并不全面:在运维安全中,每一次推理步骤、工具调用、遥测查询和情报补充都会消耗预算。我们通过这种成本-成功视角,对开放型Cybench网络攻击挑战和防御型Splunk BOTS v1调查挑战中的语言模型安全代理进行评估。不同于仅报告最佳情况下的成功率,我们对比了固定成本水平下的模型表现,并按推理开销和工具开销对性能进行分解。研究结果揭示了红蓝队任务中不同的扩展规律:进攻型CTF性能随测试时计算量的增加而提升,且扩展后的开源权重模型在保持成本竞争力的情况下可接近前沿商业系统;而防御型SOC调查则不具备相同的扩展特性——其成功更多依赖于工具使用的规范性、遥测数据的导航能力以及情报的精准筛选,而非单纯的推理预算。我们认为,安全代理基准测试应将经济效率和运维适配度与任务成功率并重考量。融入成本意识的SOC原生评估能更清晰地揭示哪些模型在当前实际可用,以及防御型代理仍需在哪些环节改进。我们已将研究结果发布在交互网站 https://evals.frontier.security。
English
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.