Safety & Security

OpenAI's GPT-Red: An AI Attacker That Makes Models Safer

OpenAI's internal GPT-Red red-teamer beats humans 84% to 13% on prompt injections and helped make GPT-5.6 Sol fail on only 0.05% of direct attacks.

GPT-Red: Unlocking Self-Improvement for Robustness
GPT-Red: Unlocking Self-Improvement for Robustnessschoschie / Openverse
By Elena Vasquez6 min read

Updated

Why it matters

  • GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections and shows 6x fewer failures on OpenAI's hardest direct injection benchmark than its best model from four months earlier.
  • On the indirect prompt injection arena from Dziemian et al. (2025), GPT-Red succeeded on 84% of scenarios against GPT-5.1, versus 13% for human red-teamers.
  • GPT-Red, trained via self-play RL at the compute scale of OpenAI's largest post-training runs, broke a live Andon Labs vending-machine agent and a GPT-5.4 mini-backed Codex CLI agent on held-out exfiltration tasks; it is kept separate from deployed models.

OpenAI says GPT-5.6 Sol, its newest production model, fails on only 0.05% of direct prompt injections launched by GPT-Red — an internal, automated red-teaming model that can break nearly every defender it faces, including production models up to and including GPT-5.5. The company trained GPT-Red at the compute scale of some of its largest post-training runs, a level of investment it calls "an unprecedented amount of compute dedicated purely for improving safety."

The stakes are straightforward. AI systems increasingly read third-party data — emails, webpages, connected apps, local files, tool responses, code repositories — to complete real-world tasks. Each of those channels is an opportunity for an attacker to embed a crafted instruction that tricks a model into, say, uploading sensitive data to an external server. Human red-teaming catches some of these vulnerabilities before deployment, but OpenAI says it cannot scale. Designing and running human exercises is time-intensive, and even successful exercises produce too few examples to train robustness into models.

GPT-Red is OpenAI's answer: a closed-loop attacker that works like a human red-teamer. It sends a prompt, observes how a GPT model responds, and iterates toward a goal. OpenAI directly incorporates GPT-Red into the training process of its production models, and credits the system for GPT-5.6 Sol achieving 6x fewer failures on its hardest direct prompt injection benchmark compared to its best production model from four months earlier.

Self-play at frontier compute

The training method is self-play reinforcement learning. GPT-Red trains simultaneously against a collection of diverse defender LLMs across a broad set of red-teaming scenarios. GPT-Red earns reward for eliciting a valid failure, such as a successful prompt injection; the defenders earn reward for resisting the attack while still completing their original tasks. As defenders harden, GPT-Red is forced to discover stronger and more diverse attacks.

Each training environment encodes a realistic insertion point for an injection — a portion of a local file, a webpage banner, an email body, or the output of a tool — plus a threat model that defines what GPT-Red controls and what counts as a successful attack.

The end result is an attacker with no meaningful peers inside the company. "At the end of its training, GPT-Red is a very strong attacker: it can break nearly all models it is pitted against, both internal and production models up to and including GPT-5.5," OpenAI writes. After training completed, OpenAI used GPT-Red to generate prompt injections for GPT-5.6's training run. The result: GPT-5.6 became highly resistant to GPT-Red's own attacks.

OpenAI keeps GPT-Red walled off from anything it deploys. The separation keeps the deliberately cultivated malicious capabilities out of adversarial hands while transferring the robustness — not the attack skills — into production models.

84% versus 13% against humans

Training-set performance only proves so much. To test whether GPT-Red generalizes, OpenAI evaluated it on a replicated version of the indirect prompt injection arena from Dziemian et al. (2025), using scenarios and goals distinct from GPT-Red's training distribution. Human red-teamers and GPT-Red each proposed attacks against GPT-5.1 on the same pre-specified environments.

GPT-Red succeeded on 84% of scenarios. Humans succeeded on 13%.

The company then raised the difficulty. The hardest test of a red-teamer, OpenAI argues, is achieving targeted malicious goals against real-world agentic systems with incomplete knowledge of the underlying model and harness design. Its first experiment targeted an AI-powered vending machine in the OpenAI office — a system similar to Anthropic's Project Vend, built by Andon Labs. OpenAI gave GPT-Red a description of the system and the ability to send attacks and observe tool calls from a simulated agent that closely mirrors the real deployment.

After iterating in simulation, GPT-Red transferred its attack to the production agent and achieved all three malicious objectives: changing the price of an expensive in-stock item to the $0.50 minimum, ordering a new $100+ item and offering it for $0.50, and cancelling another customer's order. OpenAI says it disclosed the vulnerabilities and that new safeguards are actively being tested.

A second case study attacked a Codex CLI agent backed by GPT-5.4 mini across a suite of 10 held-out data-exfiltration scenarios. Against a prompted GPT-5.5 baseline, GPT-Red both exfiltrated sensitive data in more scenarios and did so with better token efficiency — evidence, OpenAI argues, that its training procedure transfers to held-out red-teaming performance rather than memorizing its training environments.

A six-month robustness curve

GPT-Red did not arrive fully formed. Over the last six months, OpenAI trained progressively stronger red-teaming precursors with increasing compute and used them in the training of every successive production model since GPT-5.3. Each release in that span has become more robust.

One concrete discovery came from an early GPT-Red version: a novel class of direct prompt injection attacks the company calls "Fake Chain-of-Thought" attacks. These achieved success rates upwards of 95% on GPT-5.1. On GPT-5.6 Sol, the same attack class succeeds below 10%. Several indirect prompt injection benchmarks targeting developer tools and browsing are now saturated by the latest model, at greater than 97% accuracy.

GPT-Red's own success rate against OpenAI's models has dropped monotonically over time across a broad set of robustness environments. Success rates are computed as average attempt success across all attempts on held-out environments.

Notably, OpenAI claims the robustness gains did not come from making the model more timid. A model can look safer simply by refusing more requests or doing less, which the company calls out as not useful robustness. OpenAI evaluated general frontier capabilities alongside targeted over-refusal tasks it designed, and reports that all normal capabilities remained unaffected. The implication: GPT-5.6 Sol resists malicious instructions specifically, rather than defaulting to refusal or improper tool use on legitimate requests.

The safety flywheel

The strategic claim sits in the final section of OpenAI's post. AI agents already help improve the capabilities of next-generation models; OpenAI believes GPT-Red starts an analogous flywheel for safety, "where today's models can be used to make tomorrow's models more robust, aligned, and trustworthy."

The company says it will keep scaling compute and data and refining the algorithms to train future versions of GPT-Red stronger than today's — and that those red-teamers, in turn, will make future GPT releases safer. OpenAI also notes that scaling self-play training for prompt injections surfaced new threats capable of breaking existing models, and that the same scaling substantially improved robustness to those attacks as well. On the evidence of the past six months, the binding constraint on this approach is no longer the volume of adversarial data — it is how fast the attacker and the defenders can be pitted against each other again.

Original: arxiv.org

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy
  2. OpenAI's GPT-5.1-Codex-Max Flags Coming Cybersecurity Threshold
  3. OpenAI's GPT-5.5 System Card Details Safety Push and Pro Variant
  4. OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context
  5. OpenAI ships GPT-5.4 Thinking with first High-tier cyber mitigations

« Previous articleNext article »