ToolHazard: LLM 기반 에이전트의 보안 평가 및 정렬을 위한 적대적 환경 확장
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
August 12, 2026
저자: Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye
cs.AI
초록
외부 도구와 통합된 대규모 언어 모델(LLM) 에이전트는 환경 상태에 삽입된 간접 프롬프트 인젝션에 취약하다. 그러나 기존 연구는 주로 수동으로 구현되거나 재사용된 환경, 확률적 LLM 기반 도구 시뮬레이션, 사전 정의된 인젝션 위치에 의존하여 더 넓은 도메인에 걸친 확장 가능한 보안 연구를 제한한다. 이러한 격차를 해소하기 위해 우리는 인간의 엔지니어링을 줄이고 추가 시드 도메인과 컴퓨팅으로 확장을 지원하는 확장 가능한 적대적 환경 합성 프레임워크인 **ToolHazard**를 제안한다. ToolHazard는 환경 시뮬레이터, 공격자 에이전트, 사용자 시뮬레이터를 통해 실행 가능한 상태 저장 환경을 합성하고, 유효한 인젝션 지점을 발견하며 환경별 페이로드를 생성하고, 상태 기반 장기 과제를 구축한다. ToolHazard를 기반으로 복잡한 워크플로우와 다양한 환경 공격 하에서 에이전트를 스트레스 테스트하기 위한 **ToolHazard-Bench**를 구축한다. 실험은 상당한 에이전트 취약점을 드러내며, 인젝션 시점과 위치가 공격 효과에 영향을 미침을 보여준다. 또한, ToolHazard가 생성한 정렬 데이터는 ToolHazard-Bench와 AgentDojo 모두에서 보안을 개선하면서 정상 작업 유용성을 보존한다.
English
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.