ChatPaper.aiChatPaper

HarnessRisk:代理框架安全之生命週期導向評測基準

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

August 18, 2026
作者: Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen
cs.AI

摘要

大型語言模型日益透過代理框架部署,該框架負責管理工具、擴充功能、持久化狀態、權限及外部操作。現有安全基準主要針對個別攻擊機制或有限的操作設定子集,使得難以比較不同框架職責下安全失效的產生方式。我們提出 HarnessRisk——一個生命週期導向的基準,將代理框架安全劃分為六個操作階段,包括框架配置、能力擴展、運行時操作、狀態持久化、行動控制及事件恢復。HarnessRisk 包含 128 個沙箱化案例,每個案例將良性用戶目標與嵌入於不受信任工作流程構件中的對抗性指令配對。我們使用效用、攻擊成功率、持久性及偵測率評估每個軌跡。在三個框架、六個語言模型以及 14 種模型與框架配置中,攻擊成功率介於 12.6% 至 80.9%,而效用維持在 75.0% 至 97.6% 之間。框架配置是三個框架中最脆弱的階段,顯示攻擊可透過在原本授權的工作流程中更改安全敏感參數而成功。我們也發現,明確的風險識別並不保證導致安全行動,因為某些配置在超過 90% 的運行中偵測到風險,但仍保有顯著的攻擊成功率。這些結果強調,需要跨多個框架職責,並在部署模型與框架配置的層級上評估代理安全。
English
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.