ChatPaper.aiChatPaper

GPT-Red: 대규모 셀프 플레이를 통한 자동화된 레드 팀 공격

GPT-Red: Automated Red Teaming via Self-Play at Scale

July 28, 2026
저자: Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
cs.AI

초록

저희는 최첨단 LLM을 대상으로 새로운 프롬프트 인젝션 공격을 발견하도록 훈련된 자동화된 레드티밍 에이전트인 GPT-Red를 소개합니다. 이 모델의 목표는 프로덕션 시스템의 견고성을 평가하고 개선하는 것입니다. 이를 위해 저희는 이 모델을 활용하여 현재까지 프롬프트 인젝션에 가장 강력한 모델인 GPT-5.6을 적대적으로 훈련합니다. GPT-Red를 만들기 위해, 동시에 훈련되는 다양한 방어자 에이전트 집단을 공격하는 과제를 모델에 부여하는 확장 가능한 자기 대체(self-play) 알고리즘을 설계했습니다. 저희는 가장 큰 규모의 RL 사후 훈련 실행과 동일한 수준의 컴퓨팅 자원을 사용하여 현실적인 레드티밍 환경에서 모델을 훈련했으며, 이는 지금까지 문서화된 LLM 안전 훈련 중 가장 큰 규모입니다. GPT-Red는 레드티밍에 탁월합니다: GPT-5.5까지의 이전 모델을 안정적으로 돌파하며, 인간 레드티머보다 더 많은 성공적인 공격을 발견하고, 보류된 환경, 방어자 모델, 하네스(harnesess)에도 일반화됩니다. 향후에는 각 새로운 GPT 모델의 견고성이 개선됨에 따라, 이는 더 강력한 레드티밍 에이전트를 위한 더 나은 학습 신호를 제공하여 자기 개선 선순환을 가능하게 할 것으로 기대합니다.
English
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.