GPT-Red:大规模自我博弈驱动的自动化红队测试
GPT-Red: Automated Red Teaming via Self-Play at Scale
July 28, 2026
作者: Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
cs.AI
摘要
我们介绍GPT-Red,这是一种自动化的红队智能体,专门训练用于发现针对前沿大语言模型的新型提示注入攻击。该模型的目标是评估并增强生产系统的鲁棒性。为此,我们利用它对GPT-5.6进行对抗性训练,这是迄今为止对提示注入最鲁棒的模型。为构建GPT-Red,我们设计了一种可扩展的自对弈算法,要求模型攻击一组同时训练的多样化防御者智能体。我们在逼真的红队环境中训练该模型,所使用的计算规模与部分最大的强化学习后训练运行相当,使其成为有史以来文档记录中最大的LLM安全训练项目。GPT-Red在红队测试中表现出色:它能可靠地攻破我们此前直至GPT-5.5的模型,发现比人类红队成员更多的成功攻击,并且能泛化到未见过的环境、防御者模型及工具链。未来,我们预期随着每个新GPT模型鲁棒性的提升,它将反过来为更强大的红队智能体提供更优的学习信号,从而开启自我改进的飞轮效应。
English
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.