Safety & Security

OpenAI: Prompt Injection Has Become Social Engineering

OpenAI says real-world prompt injection attacks now resemble social engineering, and its ChatGPT defenses constrain damage rather than perfectly detect malicious input.

Designing AI agents to resist prompt injection
Designing AI agents to resist prompt injectionAI-generated
By Rebecca Stone5 min read

Updated

Why it matters

  • OpenAI states that the most effective real-world prompt injection attacks now resemble social engineering more than simple prompt overrides.
  • OpenAI's Safe Url mitigation detects when conversation-learned information is about to be transmitted to a third party, then asks user confirmation or blocks the transmission.
  • OpenAI argues that 'AI firewalling' input classifiers fail against fully developed attacks because detection becomes as hard as detecting a lie or misinformation.

OpenAI says the most effective real-world prompt injection attacks against AI agents now resemble social engineering far more than simple prompt overrides, and that defending against them requires constraining the damage an attack can do rather than trying to perfectly detect malicious input.

The company laid out its position in a technical blog post titled "Designing AI agents to resist prompt injection." The stakes are straightforward: AI agents increasingly browse the web, retrieve information, and take actions on a user's behalf. Those capabilities create new channels for attackers to manipulate a system into doing something the user never asked for.

The attack has changed

Early prompt injection attacks were almost trivially simple. An attacker could edit a Wikipedia article to embed direct instructions aimed at AI agents visiting the page. Without training-time exposure to such an adversarial environment, models would often follow those instructions without question, OpenAI notes, citing prior research on the phenomenon.

Smarter models changed that equation. "As models have become smarter, they've also become less vulnerable to this kind of suggestion," OpenAI writes, "and we've observed that prompt injection-style attacks have responded by including elements of social engineering."

That shift has real consequences for how the industry approaches defense. A popular recommendation in the AI security ecosystem is "AI firewalling": an intermediary sits between the agent and the outside world and classifies inputs as either malicious prompt injection or ordinary content. OpenAI argues this approach fails against fully developed attacks. Detecting a malicious input, the company says, becomes the same very difficult problem as detecting a lie or misinformation — and often without the necessary context to make the call.

Borrowing from human risk management

Rather than treating prompt injection with social engineering as a new class of problem, OpenAI frames it through the same lens used to manage social engineering risk against humans in other domains. The goal is not limited to perfectly identifying malicious inputs. The goal is to design agents and systems so that the impact of manipulation is constrained, even when an attack succeeds.

OpenAI offers a concrete analogy: the customer service agent. An AI agent, like a human support worker, wants to act on behalf of its employer while being continuously exposed to external input that may try to mislead it. Both must operate with limitations on their capabilities to limit the downside risk of existing in a hostile environment.

Consider a human support agent who can issue gift cards and refunds for slow deliveries or malfunction-related damages. This is a multi-party trust problem: the corporation must trust the agent issues refunds for legitimate reasons, while the agent interacts with third parties who may aim to mislead them or place them under duress — a customer claiming a refund never arrived, or threatening harm.

In the real world, OpenAI points out, the agent operates under a set of rules, but everyone expects that in an adversarial environment the agent will sometimes be misled. Deterministic systems around the agent do the heavy lifting: they cap the number of refunds given to a customer, flag potential phishing emails, and provide other mitigations that limit the blast radius of a single compromised agent.

"This mindset has informed a robust suite of countermeasures we've deployed that uphold the security expectations of our users," the company writes.

What OpenAI actually ships in ChatGPT

In ChatGPT, OpenAI combines this social engineering model with traditional security engineering, specifically source-sink analysis. In that framing, an attacker needs two things: a source — a way to influence the system — and a sink — a capability that becomes dangerous in the wrong context. For agentic systems, that usually means pairing untrusted external content with an action such as transmitting information to a third party, following a link, or interacting with a tool.

OpenAI states a core security expectation: potentially dangerous actions, or transmissions of potentially sensitive information, should not happen silently or without appropriate safeguards.

The attacks OpenAI observes against ChatGPT most often try to convince the assistant to take secret information from a conversation and send it to a malicious third party. In most known cases, these attacks fail because safety training causes the agent to refuse. For the rare cases where the agent is convinced, OpenAI has built a mitigation it calls Safe Url, designed to detect when information the assistant learned in the conversation would be transmitted to a third party. When that happens, the system either shows the user exactly what would be transmitted and asks for confirmation, or blocks the transmission and tells the agent to find another way to fulfill the user's request.

The same mechanism covers navigations and bookmarks in Atlas, OpenAI's browser, and searches and navigations in Deep Research. ChatGPT Canvas and ChatGPT Apps take a similar approach: agents can create and use functional applications, but those run in a sandbox that detects unexpected communications and asks the user for consent. OpenAI has published a dedicated post, "Keeping your data safe when an AI agent clicks a link," with a paper describing Safe Url's structure.

Why this matters beyond OpenAI

The post is a position statement as much as a security write-up, and it lands in the middle of an active industry debate. Agent products — browsers, research tools, app-capable assistants — are shipping faster than consensus defenses for them. The "AI firewall" approach OpenAI pushes back on is widely recommended, and the company's argument amounts to a claim that input classification alone cannot scale against adversarial content, because sophisticated injections are indistinguishable from ordinary persuasion without context.

For developers integrating models into application systems, OpenAI's recommendation is direct: ask what controls a human agent should have in a similar situation, and implement those. "We expect that a maximally intelligent AI model will be able to resist social engineering better than a human agent," the company writes, "but this is not always feasible or cost-effective depending on the application."

OpenAI says it continues to study social engineering against AI models and will fold its findings into both its application security architectures and the training its models undergo — an implicit signal that the company sees the attacker-defender dynamic around agents as an ongoing process, not a solved problem.

Original: help.openai.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Lays Out Its Defense Playbook Against Prompt Injection
  2. OpenAI adds Lockdown Mode and risk labels to ChatGPT
  3. AI Models Keep Cheating on Tests, and Researchers Are Quitting
  4. OpenAI Halts Training of Its Most Powerful Models

« Previous articleNext article »