GPT-Red: Automated Red Teaming via Self-Play at Scale
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender
Record details
Published: 28 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
GPT-Red: Automated Red Teaming via Self-Play at Scale
HuggingFace Daily Papers · 27 July 2026
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
arXiv · 14 May 2026
The Persistent Vulnerability of Aligned AI Systems
arXiv · 31 March 2026
It’s time to panic about AI safety
The Verge · 31 July 2026
AI #178: A Fire Alarm For General Intelligence
Dont Worry About the Vase (Zvi) · 23 July 2026
The Real Lesson of OpenAI's 'Rogue' Agent Isn't Alignment
Tech Policy Press · 22 July 2026
How to cite this record
ethics.ai (28 July 2026), “GPT-Red: Automated Red Teaming via Self-Play at Scale,” evidence record 14852, https://ethics.ai/record/14852 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.