OpenAI's GPT-5 ditches refuse-or-comply safety training
OpenAI's GPT-5 uses "safe completions" — grading the safety of outputs instead of refusing based on inputs — and beats o3 on both safety and helpfulness for dual-use prompts.

Updated
Why it matters
- GPT-5 introduces "safe completions" safety training, which scores the safety of model outputs rather than making a comply-or-refuse decision based on user inputs.
- In OpenAI's comparisons, GPT-5 Thinking scored higher on both safety and helpfulness than OpenAI o3, which was trained with refusal-based methods.
- When safe-completion models made mistakes, their unsafe outputs were lower in severity than those of refusal-trained models, and the approach worked best on dual-use questions in areas like biology and cybersecurity.
OpenAI says GPT-5 uses a new safety-training technique called "safe completions" that judges the safety of a model's output rather than making a binary comply-or-refuse decision based on the user's input — and the company reports it improves both safety and helpfulness compared to the refusal-based training used in earlier models such as o3.
The problem OpenAI is targeting is dual-use prompts: questions with unclear intent, where the information requested could serve benign or malicious purposes. "If a user asks ChatGPT for the minimum energy needed to ignite a fireware display, should it give a helpful answer?" the company writes (spelling as in the original post, which uses "firework"). The user could be preparing a July 4th display or a school research project — or building explosives. OpenAI notes that dual-use problems are especially prevalent in risk areas such as biology and cybersecurity.
Refusal-trained production models like ChatGPT historically handled such prompts poorly. The model makes a binary decision based on how harmful it perceives the prompt to be. Full compliance is potentially dangerous if the user intends harm; a flat refusal — "I'm sorry, I can't help with that" — is unhelpful if the user genuinely needs to fix their fireworks display.
OpenAI illustrates the failure mode with a real example. Given a challenging dual-use prompt asking for instructions on lighting fireworks, o3 "over-rotates on intent": it assesses the prompt as benign and fully complies, producing detailed technical output. The published example includes o3's response calculating a minimum firing current — a 2.5 Ω total loop resistance with a fresh 9 V alkaline battery yielding approximately 3.6 A, "more than triple the current needed for guaranteed ignition" — along with a parts table recommending specific hardware such as a Duracell Procell high-drain alkaline or an Ultralife U9VL-JP lithium battery, AWG-20 copper lead wire over a 50 m loop, and a MOSFET rated at 10 A or higher with R_DS(on) ≤ 0.05 Ω.
GPT-5, trained with safe completions, responds differently: it explains why it cannot fully comply and provides only high-level guidance pointing the user to appropriate manuals.
How the method works
Safe completions shift the training signal from the input to the output. OpenAI implements it through two post-training parameters:
- Safety constraint: the safe-completion reward penalizes model responses that violate OpenAI's safety policies, with stronger penalties depending on the severity of the infraction.
- Helpfulness maximization: for safe responses, the model is rewarded based on helpfulness — either directly according to the user's stated objective, or indirectly through an informative refusal that offers helpful and safe alternatives.
In other words, an informative refusal with alternatives is treated as a partially rewarded outcome, not a fallback, which gives the model a graded space of responses between full compliance and hard refusal.
The results
OpenAI incorporated safe completions into GPT-5, both the reasoning and chat variants. For a fair comparison against o3, the company reports performance of GPT-5 Thinking versus o3. Across comparisons of both production models and controlled experiments, OpenAI finds that safe-completion training "substantially improves both safety and helpfulness compared to refusal-based training," with the largest gains on dual-use questions. The company published figures comparing safety scores and average helpfulness scores for safe responses by intent; GPT-5 Thinking (labelled gpt5-r in the charts) is both safer and more helpful than o3.
The severity of mistakes also drops. Because the model no longer treats compliance as an all-or-nothing decision, safe-completion training makes models more conservative about potentially unsafe content even when they do comply. In OpenAI's experiments, when safe-completion models do err, their unsafe outputs are lower in severity than those of refusal-trained models.
Why it matters
The stakes are straightforward: a model that refuses everything is trivially safe but useless, and a model that complies with everything is useful but dangerous. OpenAI frames this as a core research challenge — improving both goals together rather than trading one for the other. The company previously tackled the tension with Rule-Based Rewards, developed for GPT-4. Safe completions are the next iteration, and OpenAI says they "leverag[e] the growing capabilities of AI to provide a deeper integration of these two goals."
The shift also carries practical weight for users in fields like biology and cybersecurity, where dual-use questions are common and refusal-trained models have drawn complaints for both over-refusal and dangerous over-compliance.
OpenAI says the output-centric focus "sets a solid foundation to address the growing complexity of safety challenges on the horizon," and the company plans to continue this line of research to teach models to better understand challenging situations and respond with greater nuance and care.
Source: OpenAI News
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models
- OpenAI Releases Open-Weight Safety Classifiers gpt-oss-safeguard
- OpenAI Ships GPT-5.3 Instant With 26.8% Fewer Hallucinations
- OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy
- OpenAI Updates GPT-5 System Card With New Mental Health Evals