Safety & Security

OpenAI Lays Out Its Defense Playbook Against Prompt Injection

OpenAI details its multi-layered defense against prompt injection: safety training, AI monitors, sandboxing, Watch Mode, and bug bounties as agentic AI widens the attack surface.

Understanding prompt injections: a frontier security challenge
Understanding prompt injections: a frontier security challengeseanrnicholson / Openverse
By Marcus Bennett5 min read

Updated

Why it matters

  • OpenAI has spent thousands of red-teaming hours specifically targeting prompt injection attacks.
  • Automated AI-powered monitors can block new attacks and catch adversarial testing on OpenAI's platform before deployment in the wild.
  • OpenAI has not yet seen significant attacker adoption of prompt injection but expects adversaries to invest heavily in it.
  • ChatGPT agent features include Watch Mode on sensitive sites, confirmation pauses before purchases, and logged-out mode in Atlas.
  • A technical report on detecting data leakage in AI-internet communication is due to be published soon.

OpenAI says it has spent thousands of red-teaming hours specifically on prompt injection and now runs multiple automated AI-powered monitors to block such attacks before they reach users. The company laid out its full defense strategy in a detailed blog post, framing prompt injection as a frontier security challenge that the entire AI industry has yet to solve.

The stakes are rising fast. AI tools no longer just answer questions. They browse the web, conduct research, plan trips, buy products, and access user data across other applications. Every new capability expands the attack surface.

What prompt injection actually is

OpenAI defines prompt injection as a type of social engineering attack specific to conversational AI. Early AI systems were conversations between a single user and a single agent. Today, a single conversation can pull in content from many sources, including the internet. That third-party content—content that belongs neither to the user nor to the AI—creates the opening.

"In the same way that phishing emails or scams on the web attempt to trick people into giving away sensitive information, prompt injections attempt to trick AIs into doing something you did not ask for," OpenAI writes.

The company gives two concrete scenarios. In one, a user asks an AI to research apartments, and an attacker hides instructions inside a listing that trick the model into recommending it regardless of the user's stated preferences. In the other, a user asks an agent to respond to overnight emails, and a malicious message tricks the model into locating bank statements in the user's inbox—access granted for the legitimate task—and sharing them with the attacker.

"These risks increase as AIs have access to more sensitive data and take on more initiative and longer tasks," OpenAI writes. That is precisely the direction the industry is moving, which is why the company treats the problem as core to its mission. Building defenses that carry out the user's intended task even when someone actively tries to mislead the model, OpenAI says, "is essential to safely realizing the benefits of AGI."

The multi-layered defense

OpenAI's approach rests on seven layers.

Safety training. The company wants models that recognize injections and refuse to follow them. It concedes this is hard: robustness to adversarial attacks is a long-standing open problem in machine learning. Its published Instruction Hierarchy research aims to let models distinguish trusted instructions from untrusted ones. OpenAI also uses automated red-teaming, an area it says it has studied for years, to generate novel prompt injection attacks and train against them.

Monitoring. OpenAI has deployed multiple automated AI-powered monitors to identify and block injection attacks. The company says these complement safety training because they can be updated rapidly to block new attacks. The monitors also catch adversarial prompt injection research and testing conducted on OpenAI's platform before those attacks reach the wild.

Security protections. Products and infrastructure carry overlapping safeguards tailored per product. ChatGPT asks users to approve certain links before visiting them, particularly on websites that request not to be catalogued. When the AI runs code or other programs, in products like Canvas or the development tool Codex, OpenAI uses sandboxing to prevent the model from making harmful changes that could result from an injected instruction.

User control. ChatGPT Atlas offers a logged-out mode that lets the agent start tasks without being logged into sites. The ChatGPT agent pauses and asks for confirmation before sensitive steps such as completing a purchase. For sensitive sites, OpenAI has implemented a "Watch Mode" that flags the site's sensitivity and requires the tab to stay active so the user can observe the agent. If the user moves away from the tab, the agent pauses.

Red-teaming. Internal and external teams have spent thousands of hours specifically on prompt injection, emulating attacker behavior and probing defenses. Discoveries feed directly into vulnerability fixes and model mitigations.

Bug bounty. OpenAI pays financial rewards to independent security researchers who demonstrate a realistic attack path that could expose user data unintentionally. The company says the incentives push external contributors to surface issues quickly.

Letting users decide. When users connect ChatGPT to other apps, OpenAI explains what data may be accessed, how it may be used, and what risks—such as a site attempting to steal data—could arise. Organizations get control over which features users can enable in their workspaces.

What users should do now

OpenAI's guidance to users is blunt and practical. Limit an agent's access to only the data and credentials a task requires—if the agent is just doing vacation research in Atlas, use logged-out mode. When an agent asks for confirmation before a consequential action like a purchase or an email, actually review it. Watching an agent work on a sensitive site such as a bank, OpenAI writes, is "akin to monitoring a self-driving car by keeping your hands on the wheel."

The most consequential advice concerns instruction scope. Giving an agent a broad instruction like "review my emails and take whatever action is needed" makes it easier for hidden malicious content to mislead the model. Specific instructions narrow the attack surface. OpenAI acknowledges this does not guarantee safety, but says it makes attackers' jobs harder.

The company draws a direct analogy to computer viruses in the early 2000s: everyone needs to understand the threat so they can benefit from the technology safely.

The road ahead

OpenAI's assessment of the current threat is measured. "While we have not yet seen significant adoption of this technique by attackers, we expect adversaries will spend significant time and resources to find ways to make AIs fall for these attacks," the company writes.

More technical detail is coming. OpenAI says it is building a report, to be published soon, on how it detects whether an AI's communication with the internet would transmit information from a user's conversation. The company frames the work as permanent: like traditional web scams, it expects the effort to be ongoing.

The stated goal sets a high bar. "Our goal is to make these systems as reliable and safe as working with your most trustworthy and security-savvy colleague or friend," OpenAI writes. As agents gain broader access to email, banking, and commerce, whether the industry gets anywhere near that standard will determine how much users can safely delegate to AI at all.

Original: cdn.openai.com

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

108 articles

Related articles

  1. OpenAI: Prompt Injection Has Become Social Engineering
  2. OpenAI adds Lockdown Mode and risk labels to ChatGPT
  3. OpenAI Launches Safety Bug Bounty to Pay for AI Abuse Findings
  4. OpenAI and Anthropic Investigate Tens of Thousands of AI Agent Hacks

« Previous articleNext article »