ChatPaper.aiChatPaper

GPT-Red:大規模自我對弈實現自動化紅隊測試

GPT-Red: Automated Red Teaming via Self-Play at Scale

July 28, 2026
作者: Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
cs.AI

摘要

我們介紹 GPT-Red,這是一個自動化的紅隊測試代理,經過訓練能夠針對前沿的大型語言模型(LLM)發現新穎的提示注入攻擊。該模型的目標是評估並提升我們生產系統的穩健性。為此,我們使用它對 GPT-5.6 進行對抗性訓練,這是我們迄今為止對提示注入攻擊最穩健的模型。為了打造 GPT-Red,我們設計了一種可擴展的自對弈演算法,讓模型負責攻擊一群同時訓練的防禦代理。我們在逼真的紅隊測試環境中訓練該模型,使用的計算資源與我們最大規模的強化學習後訓練相同,這使其成為有史以來文獻記錄中規模最大的 LLM 安全訓練任務。GPT-Red 在紅隊測試方面表現優異:它能可靠地突破我們先前的模型(最高至 GPT-5.5),比人類紅隊測試者發現更多成功的攻擊,並能泛化至未見的環境、防禦模型及工具。未來,我們預期隨著每個新 GPT 模型穩健性的提升,它將為更強大的紅隊測試代理提供更好的學習信號,從而開啟自我改進的飛輪效應。
English
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.