ChatPaper.aiChatPaper

HarnessRisk: 에이전트 하네스 안전성을 위한 생애주기 중심 벤치마크

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

August 18, 2026
저자: Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen
cs.AI

초록

대규모 언어 모델은 도구, 확장 기능, 지속 상태, 권한 및 외부 작업을 관리하는 에이전트 하네스를 통해 점점 더 많이 배포되고 있다. 기존 안전성 벤치마크는 주로 개별 공격 메커니즘 또는 제한된 운영 설정의 하위 집합을 대상으로 하므로, 서로 다른 하네스 책임 전반에서 안전 실패가 어떻게 발생하는지 비교하기 어렵다. 우리는 에이전트 하네스 안전을 하네스 구성, 기능 확장, 런타임 운영, 상태 지속성, 동작 제어 및 사고 복구를 포함한 여섯 가지 운영 단계로 구성하는 수명주기 기반 벤치마크인 HarnessRisk를 제시한다. HarnessRisk는 128개의 샌드박스 케이스를 포함하며, 각 케이스는 정상적인 사용자 목표와 신뢰할 수 없는 워크플로 아티팩트에 내장된 적대적 지시를 짝지은 것이다. 우리는 각 궤적을 유틸리티, 공격 성공률, 지속성 및 탐지를 사용하여 평가한다. 세 개의 하네스, 여섯 개의 언어 모델, 그리고 14개의 모델 및 하네스 구성에서 공격 성공률은 12.6%에서 80.9%까지 다양했으며, 유틸리티는 75.0%에서 97.6% 사이를 유지했다. 하네스 구성은 세 하네스 모두에서 가장 취약한 단계로, 공격이 정상적으로 승인된 워크플로 내에서 보안에 민감한 매개변수를 변경함으로써 성공할 수 있음을 보여준다. 또한 일부 구성은 전체 실행의 90% 이상에서 위험을 탐지하면서도 상당한 공격 성공률을 유지하므로, 명시적 위험 인식이 안전한 조치로 안정적으로 이어지지 않음을 발견한다. 이 결과는 여러 하네스 책임 전반과 배포된 모델 및 하네스 구성 수준에서 에이전트 안전성을 평가해야 할 필요성을 강조한다.
English
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.