GPT-Red: 大規模な自己対戦による自動レッドチーミング
GPT-Red: Automated Red Teaming via Self-Play at Scale
July 28, 2026
著者: Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
cs.AI
要旨
本稿では、最先端LLMに対する新たなプロンプトインジェクション攻撃を発見するために訓練された自動レッドチーミングエージェント「GPT-Red」を紹介する。本モデルの目的は、本番システムの堅牢性を評価・改善することである。そのために、我々はGPT-Redを用いて、現時点でプロンプトインジェクションに対して最も堅牢なモデルであるGPT-5.6を敵対的訓練する。GPT-Redを構築するにあたり、同時に訓練される多様な防御エージェント群に対して攻撃を仕掛けるタスクをモデルに課す、スケーラブルな自己対戦アルゴリズムを設計した。我々は、最大規模の強化学習事後訓練実行と同等の計算資源を用いて、現実的なレッドチーミング環境でモデルを訓練した。これは、これまでに文書化された中で最大規模のLLM安全性訓練実行である。GPT-Redはレッドチーミングに優れており、過去のモデル(GPT-5.5まで)を確実に突破し、人間のレッドチーマーよりも多くの攻撃成功を発見し、未見の環境、防御モデル、ハーネスに対して汎化する。今後、各新世代GPTモデルの堅牢性が向上するにつれ、さらに強力なレッドチーミングエージェントのための学習シグナルが改善され、自己改善のフライホイールが回り始めるものと期待される。
English
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.