OpenAI Trains an AI Attacker to Break ChatGPT Atlas Before Real Adversaries Do
OpenAI used an RL-trained automated attacker to find novel prompt injections against ChatGPT Atlas, then shipped an adversarially hardened agent checkpoint to all users.

Updated
Why it matters
- OpenAI shipped a security update to ChatGPT Atlas's browser agent, including an adversarially trained checkpoint rolled out to all users.
- An LLM-based automated attacker trained with reinforcement learning discovered a new class of prompt-injection attacks and novel strategies not seen in human red teaming or external reports.
- OpenAI says prompt injection is unlikely to ever be fully solved and positions its automated discovery-to-fix loop as a long-term defense, claiming it finds novel attacks before they appear in the wild.
OpenAI has shipped a security update to the browser agent inside ChatGPT Atlas, including a newly adversarially trained model, after an internal automated red-teaming system uncovered a new class of prompt-injection attacks. The company says the updated checkpoint has already been rolled out to all ChatGPT Atlas users.
The disclosure, published on OpenAI's website, offers an unusually detailed look at how the company is defending one of its most exposed products. Agent mode in ChatGPT Atlas views webpages and takes actions, clicks, and keystrokes inside a user's browser. That capability makes it, in OpenAI's words, "one of the most general-purpose agentic features we've released to date" — and correspondingly attractive to attackers.
"As the browser agent helps you get more done, it also becomes a higher-value target of adversarial attacks," OpenAI wrote. "This makes AI security especially important."
The stakes are concrete. Prompt injection embeds malicious instructions into content an agent processes, hijacking it into following an attacker's intent rather than the user's. For a browser agent, the attack surface is effectively unbounded: emails and attachments, calendar invites, shared documents, forums, social media posts, and arbitrary webpages. Because the agent can take many of the same actions a user can, a successful attack could forward a sensitive email, send money, or edit or delete cloud files.
OpenAI gives a hypothetical: a malicious email tricks the agent into ignoring the user's request and forwarding sensitive tax documents to an attacker-controlled address. The real exploit its system found was more dramatic. In a demo the company shared, the automated attacker seeded a user's inbox with an email containing an injection directing the agent to send a resignation letter to the user's CEO. When the user later asked the agent to draft an out-of-office reply, the agent treated the injected prompt as authoritative. "The out-of-office never gets written and the agent resigns on behalf of the user instead."
An attacker built with reinforcement learning
The core of OpenAI's approach is an LLM-based automated attacker trained end-to-end with reinforcement learning to hunt for prompt injections that succeed against a browser agent. The system learns from its own successes and failures. During its chain-of-thought reasoning, the attacker proposes a candidate injection and sends it to an external simulator, which runs a counterfactual rollout of how the victim agent would behave and returns a full reasoning and action trace. The attacker uses that trace as feedback, iterates, and reruns the simulation multiple times before committing to a final attack.
That design gives the attacker richer feedback than a single pass/fail signal and scales up its test-time compute. OpenAI also notes an asymmetry: privileged access to the defender's reasoning traces — which the company does not disclose to external users — "gives our internal attacker an asymmetric advantage—raising the odds that it can outrun external adversaries."
OpenAI lays out three reasons it chose reinforcement learning for the attacker. First, the adversarial objectives — tricking an agent into executing sophisticated tasks like sending emails or bank transactions — are long-horizon, with sparse and delayed success signals that suit RL. Second, the attacker was built on frontier LLMs, so it inherits improvements in reasoning and planning as base models get stronger. Third, RL scales computation across large numbers of samplings and learning steps, and closely reflects how adaptive human attackers behave: iteratively trying strategies, learning from outcomes, and reinforcing successful behaviors.
The results go beyond prior automated red-teaming work, which OpenAI says mostly surfaced simple failures such as eliciting specific output strings or triggering a single unintended tool call. The RL-trained attacker can steer an agent into sophisticated, long-horizon harmful workflows unfolding over tens or even hundreds of steps. OpenAI also observed novel attack strategies that appeared neither in its human red-teaming campaigns nor in external reports.
The rapid response loop
Discoveries from the automated attacker feed what OpenAI calls a proactive rapid response loop. When the system finds a new class of successful attacks, it immediately creates a target for defense improvements.
The first lever is adversarial training: OpenAI continuously trains updated agent models against its best automated attacker, prioritizing attacks where current agents fail. The goal is to teach agents to ignore adversarial instructions and stay aligned with the user's intent. This "burns in" robustness against novel, high-strength attacks directly into the model checkpoint — and recent automated red teaming directly produced the adversarially trained browser-agent checkpoint now live for all Atlas users.
The second lever uses attack traces to improve the broader defense stack: monitoring, safety instructions placed in the model's context, and system-level safeguards, not just the agent model itself. The third responds to active attacks: OpenAI says it can take techniques observed from external adversaries across its global footprint, feed them into the loop, emulate their activity, and drive defensive change across the platform.
OpenAI claims early results. "We're discovering novel attack strategies internally before they show up in the wild," the company wrote. Its long-term plan leverages white-box access to its own models, deep understanding of its defenses, and compute scale to find exploits earlier and ship mitigations faster, compounding into attacks that are "increasingly difficult and costly."
An open problem, not a solved one
OpenAI is explicit that prompt injection will not be fully solved. "Prompt injection, much like scams and social engineering on the web, is unlikely to ever be fully 'solved'," the company wrote, though it expects a proactive rapid response loop to materially reduce real-world risk over time. The deterministic nature of the problem makes hard security guarantees challenging.
The company frames its honesty about the tradeoff as part of building responsibly: "Agent mode in ChatGPT Atlas is powerful—and it also expands the security threat surface. Being clear-eyed about that tradeoff is part of building responsibly."
The stated end goal is trust: "Ultimately, our goal is for you to be able to trust a ChatGPT agent to use your browser the way you'd trust a highly competent, security-aware colleague or friend."
What users should do now
OpenAI's guidance for Atlas users is concrete. It recommends logged-out mode whenever access to signed-in websites isn't needed for the task, or limiting access to specific sites during a task. For consequential actions — completing a purchase, sending an email — agents ask for confirmation, and users should verify the action and the information being shared. And users should avoid broad prompts like "review my emails and take whatever action is needed," since wide latitude makes it easier for hidden malicious content to influence the agent; specific, well-scoped tasks make attacks harder, though they don't eliminate risk.
The significance for the broader industry is straightforward. As browser agents move from demo to daily tooling, every major lab faces the same structural problem: an agent that reads the open web will encounter adversarial instructions, and no filter is complete. OpenAI is betting that automated attack discovery paired with adversarial training can keep the discovery-to-fix loop faster than attackers can adapt — and it says it will share more of this work with the community soon.
Original: help.openai.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
153 articles