ToolHazard:擴展對抗環境用於基於LLM的智能體的安全評估與對齊
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
August 12, 2026
作者: Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye
cs.AI
摘要
與外部工具整合的大型語言模型(LLM)代理容易受到嵌入在環境狀態中的間接提示注入攻擊。然而,現有研究大多依賴手動實作或重複使用的環境、基於 LLM 的隨機工具模擬,以及預先定義的注入位置,限制了在更廣泛領域中進行可擴展的安全研究。為填補此一差距,我們提出了 **ToolHazard**,一個可擴展的對抗性環境合成框架,可減少人工工程,並支援透過額外的種子領域和計算資源進行擴展。透過環境模擬器、攻擊者代理和使用者模擬器,ToolHazard 能合成可執行的有狀態環境、發現可行的注入點並產生特定環境的攻擊載荷,以及建構基於狀態的長時程任務。基於 ToolHazard,我們建立了 **ToolHazard-Bench**,用於在複雜工作流程和多樣化環境攻擊下對代理進行壓力測試。實驗揭示了代理存在重大漏洞,並顯示注入時機和位置會影響攻擊有效性。此外,ToolHazard 產生的對齊數據能提升 ToolHazard-Bench 和 AgentDojo 上的安全性,同時保持良性任務的效用。
English
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.