ChatPaper.aiChatPaper

超越成功率:成本感知評估攻擊與防禦安全智能體

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

July 17, 2026
作者: Paul Kassianik, Blaine Nelson, Yaron Singer
cs.AI

摘要

安全性代理的評估通常著眼於在充裕推理預算下的峰值攻擊能力,側重漏洞發現、漏洞利用開發、滲透測試及CTF解題。此類測量雖有助益卻不完整:在實際運營安全中,每一次推理步驟、工具調用、遙測查詢與情報補充皆會消耗預算。我們透過成本-成功視角,在攻擊型Cybench挑戰與防禦型Splunk BOTS v1調查挑戰中評估語言模型安全性代理。不同於僅報告最佳情況的成功率,我們在固定成本層級下比較各模型,並按其推理支出與工具支出拆解表現。結果顯示紅隊與藍隊任務呈現不同的擴展規律。攻擊型CTF表現隨測試階段計算量增加而提升,且擴展後的開源權重模型可接近前沿專有系統,同時保持成本競爭力。防禦型SOC調查並未以相同方式擴展:其成功更依賴於紀律性的工具使用、遙測數據導航及選擇性情報補充,而非單純的推理預算。我們主張安全代理基準應同時衡量經濟效率與運營契合度,而非僅評估任務成功。具成本意識、貼近SOC實務的評估,能更清晰呈現哪些模型今日具有實用價值,以及防禦代理仍需改進之處。我們提供互動式網站展示結果:https://evals.frontier.security
English
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.