성공률을 넘어서: 공격 및 방어 보안 에이전트의 비용 인식 평가
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
July 17, 2026
저자: Paul Kassianik, Blaine Nelson, Yaron Singer
cs.AI
초록
보안 에이전트 평가는 일반적으로 관대한 추론 예산 하에서 최대 공격 능력을 측정하며, 취약점 발견, 익스플로잇 개발, 침투 테스트, CTF 완료를 중점적으로 평가합니다. 이러한 측정은 유용하지만 불완전합니다. 실제 운영 보안에서는 모든 추론 단계, 도구 호출, 텔레메트리 쿼리, 인리치먼트 요청이 예산을 소모하기 때문입니다. 우리는 이 비용-성공 관점을 통해 공격적 Cybench 과제와 방어적 Splunk BOTS v1 조사 과제에서 언어 모델 보안 에이전트를 평가합니다. 최상의 성공 사례만 보고하는 대신, 고정된 비용 수준에서 모델을 비교하고 추론 비용과 도구 비용으로 성능을 분해합니다. 우리의 결과는 레드팀과 블루팀 과제에서 뚜렷한 확장 체계를 보여줍니다. 공격적 CTF 성능은 추가 테스트 시간 연산으로 향상되며, 확장된 오픈웨이트 모델은 비용 경쟁력을 유지하면서 최첨단 독점 시스템에 근접할 수 있습니다. 방어적 SOC 조사는 같은 방식으로 확장되지 않습니다. 성공은 순수한 추론 예산보다는 체계적인 도구 사용, 텔레메트리 탐색, 선택적 인리치먼트에 더 크게 의존합니다. 우리는 보안 에이전트 벤치마크가 과제 성공과 함께 경제적 효율성과 운영 적합성을 측정해야 한다고 주장합니다. 비용 인식적이고 SOC에 특화된 평가는 현재 실질적으로 유용한 모델이 무엇이며 방어적 에이전트가 여전히 개선해야 할 부분이 어디인지에 대한 더 명확한 그림을 제공합니다. 우리는 결과와 함께 대화형 웹사이트(https://evals.frontier.security)를 제공합니다.
English
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend. Our results show distinct scalingregimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigation does not scale in the same way: success depends more heavily on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fit alongside task success. Cost-aware, SOC-native evaluations provide a clearer picture of which models are practically useful today and where defensive agents still need to improve. We present an interactive website with our results https://evals.frontier.security.