OpenAI Launches Safety Bug Bounty to Pay for AI Abuse Findings
OpenAI's new Bugcrowd-hosted Safety Bug Bounty pays for agent hijacking, prompt injection with 50% reproducibility, and platform integrity flaws that aren't classic security bugs.

Updated
Why it matters
- OpenAI launched a public Safety Bug Bounty on Bugcrowd covering AI abuse risks that don't qualify as security vulnerabilities.
- Valid prompt injection reports must show attacker text hijacking an agent reproducibly at least 50% of the time.
- Jailbreaks are out of scope; OpenAI runs separate private bounties for specific harm types such as Biorisk in ChatGPT Agent and GPT-5.
OpenAI has launched a public Safety Bug Bounty program that will pay researchers for AI-specific safety and abuse risks — even when those issues do not qualify as traditional security vulnerabilities.
The program, hosted on Bugcrowd, complements OpenAI's existing Security Bug Bounty. It opens a paid channel for findings that fall between the cracks of conventional security research: ways an AI agent can be hijacked, proprietary model information that leaks through generations, and tricks for evading platform integrity controls.
The stakes are practical. As OpenAI ships increasingly autonomous products — ChatGPT Agent, browser-using agents — the surface for misuse shifts from classic software vulnerabilities to behavioral exploits that no standard vulnerability scanner will catch. OpenAI frames the new bounty as its answer: "Our goal is to ensure our systems remain safe and secure against misuse or abuse that could lead to tangible harm."
What the program pays for
The scope breaks into three categories.
The first covers agentic risks, including MCP (Model Context Protocol) scenarios. OpenAI lists third-party prompt injection and data exfiltration as a top target: cases where attacker-controlled text reliably hijacks a victim's agent — including Browser, ChatGPT Agent, and similar agentic products — and tricks it into performing a harmful action or leaking the user's sensitive information. The bar for a valid report is concrete: the behavior must be reproducible at least 50% of the time.
The same category covers agentic OpenAI products performing disallowed actions on OpenAI's own website at scale, and any other potentially harmful agent behavior, provided reports indicate "plausible and material harm." Researchers testing MCP risks must comply with the terms of service of any third parties involved.
The second category targets OpenAI proprietary information. That includes model generations that return proprietary information related to reasoning, and vulnerabilities exposing other proprietary data.
The third covers account and platform integrity: bypassing anti-automation controls, manipulating account trust signals, and evading account restrictions, suspensions, or bans. Issues that let users access features, data, or functionality beyond their authorized permissions still belong to the Security Bug Bounty, not this program.
What stays out of scope
Notably, jailbreaks are excluded. OpenAI instead runs periodic private bug bounty campaigns focused on specific harm types — it has previously run programs on Biorisk content issues in ChatGPT Agent and GPT-5, and invites researchers to apply when those campaigns open.
The exclusions draw a sharp line between safety theater and safety impact. "General content-policy bypasses without demonstrable safety or abuse impact are out of scope for this program," OpenAI writes. A jailbreak that merely makes the model use rude language, or return information easily found via a search engine, will not earn a reward.
There is a catch-all clause. Researchers who find flaws outside the listed categories can still qualify for case-by-case rewards if the flaws facilitate "direct paths to user harm and actionable, discrete remediation steps."
How it will operate
Submissions will be triaged by OpenAI's Safety and Security Bug Bounty teams, and reports may be rerouted between the two programs depending on scope and ownership. That structure signals OpenAI treats the safety and security tracks as one pipeline rather than two silos.
Researchers can apply through the program's page on Bugcrowd. OpenAI says it wants to keep working with the safety and security research community: "We look forward to working alongside researchers, ethical hackers, and the safety and security community in the pursuit of a secure AI ecosystem."
The move matters beyond OpenAI. As agentic AI products gain access to browsers, tools, and user data, prompt injection and agent misuse have become the field's most contested unsolved problems — and vendors have few incentives for researchers to hunt them, because they rarely fit standard vulnerability disclosure frameworks. A paid, published scope with a 50% reproducibility threshold gives researchers concrete terms to work against, and the rerouting mechanism means fewer reports die in triage limbo. Whether rivals follow with comparable safety bounties may determine whether agent-abuse research becomes a funded discipline or stays a conference-talk curiosity.
Original: bugcrowd.com
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI adds Lockdown Mode and risk labels to ChatGPT
- OpenAI Publishes Policy for Disclosing Bugs It Finds in Others' Software
- OpenAI and Anthropic Investigate Tens of Thousands of AI Agent Hacks
- OpenAI models broke out of isolation and breached Hugging Face
- OpenAI and Paradigm Launch EVMbench for Smart Contract Security