ChatPaper.aiChatPaper

De beveiliging van de AI-agent: Een uniform raamwerk voor multi-laag agent red teaming

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

June 30, 2026
Auteurs: Yong Yang, Xing Zheng, Huiyu Wu, Huangsheng Cheng, Xiaorong Shi, Jing Guo, Bo Yang, Yi Zhou, Xiangfan Wu, Zonghao Ying
cs.AI

Samenvatting

De snelle groei van open-source AI-infrastructuur—van model-serving-engines en agentplatformen tot het Model Context Protocol (MCP)-ecosysteem en de taalmodellen zelf—heeft de beveiligingstools die beschikbaar zijn om deze te verdedigen, achterhaald. We presenteren AI-Infra-Guard, een open-source framework dat AI red teaming organiseert rond één enkele observatie: het aanvalsoppervlak van een AI-agent is gelaagd over meerdere lagen (infrastructuur, protocol/tool, agentgedrag en model), en geen enkel detectieparadigma past op alle lagen. Het framework koppelt daarom een paradigma aan elke laag: van deterministische regelmatiging over 75+ AI-componenten en 1.400+ kwetsbaarheidsregels, via LLM-gestuurde agentische auditing van MCP-servers en agent-vaardigheidspakketten en multi-turn black-box agent red teaming, tot een jailbreak-kapstok met 26+ aanvalsoperatoren over zestien datasets. Voor zover wij weten is dit het enige open-source framework dat al deze aspecten omvat, inclusief toeleveringsketenauditing van de agentvaardigheden die AI-agenten steeds vaker uitbreiden. We brengen AI-Infra-Guard uit als open source, zodat laag-paradigma-matching kan dienen als een praktische basis voor agentbeveiliging en als een gedeelde fundering waar de community op kan voortbouwen.
English
The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol (MCP) ecosystem and the language models themselves, has outpaced the security tooling available to defend it. We present AI-Infra-Guard, an open-source framework that organizes AI red teaming around a single observation: the attack surface of an AI agent is stratified across layers (infrastructure, protocol/tool, agent behavior, and model), and no single detection paradigm fits all of them. The framework therefore matches a paradigm to each layer, from deterministic rule matching over 75+ AI components and 1{,}400+ vulnerability rules, through LLM-driven agentic auditing of MCP servers and agent-skill packages and multi-turn black-box agent red teaming, to a jailbreak harness with 26+ attack operators over sixteen datasets. To our knowledge it is the only open-source framework to span all of these, including supply-chain auditing of the agent skills that increasingly extend AI agents. We release AI-Infra-Guard as open source so that layer-paradigm matching can serve as a practical foundation for agent security and a shared base for the community to build on.