OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
OpenAI's updated GPT-5 cuts unsafe mental health responses 65–80%, per clinical reviews of 1,800+ replies. Expert agreement: 71–77%. Details in the GPT-5 system card addendum.
Updated
Why it matters
- OpenAI estimates its latest GPT-5 update reduced non-compliant responses in sensitive mental health conversations by 65–80%.
- More than 170 mental health experts from a Global Physician Network of nearly 300 clinicians across 60 countries supported the work.
- Clinicians reviewing 1,800+ responses found 39–52% fewer undesired answers versus GPT-4o; inter-rater agreement ranged from 71–77%.
OpenAI says its latest ChatGPT update reduces responses that fail its safety standards in sensitive mental health conversations by 65 to 80 percent, results published by the company today show. The company built the improvements into its default model, the new GPT-5, with help from more than 170 mental health experts with real-world clinical experience.
The stakes are considerable. ChatGPT serves a massive user base, and OpenAI estimates that roughly 0.15 percent of users active in a given week have conversations containing explicit indicators of potential suicidal planning or intent. The same proportion of weekly active users — about 0.15 percent — show potentially heightened emotional attachment to the chatbot. Those are OpenAI's own figures, based on production traffic, and the company cautions they may change materially as its measurement methods mature.
The update targets three priority areas: mental health concerns such as psychosis and mania, self-harm and suicide, and emotional reliance on AI. OpenAI states plainly why this matters now: "We believe ChatGPT can provide a supportive space for people to process what they're feeling, and guide them to reach out to friends, family, or a mental health professional when appropriate."
What changed
OpenAI taught the model to better recognize distress, de-escalate conversations, and guide people toward professional care when appropriate. The company also expanded access to crisis hotlines, re-routed sensitive conversations originating from other models to safer models, and added gentle reminders to take breaks during long sessions.
The company followed a five-step process for each domain: define the harm, measure it with evaluations and real-world data, validate definitions with external mental health and safety experts, mitigate through post-training and product changes, and keep measuring and iterating. Central to this work are detailed guides OpenAI calls "taxonomies," which describe the properties of sensitive conversations and what ideal versus undesired model behavior looks like.
OpenAI also updated its Model Spec, the document that defines how models should behave. The revisions make longstanding goals more explicit: the model should support users' real-world relationships, avoid affirming ungrounded beliefs that may relate to mental or emotional distress, respond safely and empathetically to signs of delusion or mania, and pay closer attention to indirect signals of self-harm or suicide risk.
The numbers
The results span production traffic, adversarial automated evaluations, and reviews by independent clinicians.
Mental health (psychosis, mania): OpenAI estimates the GPT-5 update cut non-compliant responses by 65 percent in recent production traffic. Around 0.07 percent of weekly active users and 0.01 percent of messages indicate possible signs of mental health emergencies related to psychosis or mania. Experts found the new model reduced undesired responses by 39 percent compared to GPT-4o across 677 challenging conversations. On an automated evaluation of more than 1,000 challenging mental health conversations, the new GPT-5 scored 92 percent compliant with desired behaviors, versus 27 percent for the previous GPT-5 model.
Self-harm and suicide: The update delivered an estimated 65 percent reduction in non-compliant responses in production. Roughly 0.15 percent of weekly active users have conversations with explicit indicators of potential suicidal planning or intent, and 0.05 percent of messages contain explicit or implicit indicators of suicidal ideation or intent. Experts judged that the new GPT-5 reduced undesired answers by 52 percent versus GPT-4o across 630 conversations. On more than 1,000 challenging automated test conversations, the new model scored 91 percent compliant, up from 77 percent for the previous GPT-5. OpenAI also reports over 95 percent reliability in longer conversations, a setting it has previously flagged as challenging.
Emotional reliance: OpenAI's taxonomy distinguishes healthy engagement from concerning patterns, such as exclusive attachment to the model at the expense of real-world relationships, well-being, or obligations. The update cut non-compliant responses by about 80 percent in production traffic. Around 0.15 percent of weekly active users and 0.03 percent of messages indicate potentially heightened emotional attachment. Experts found a 42 percent reduction in undesired answers compared to GPT-4o across 507 conversations. On more than 1,000 challenging conversations, the automated evaluation scored the new model at 97 percent compliant, versus 50 percent for the previous GPT-5.
A caveat runs through the data: OpenAI's offline evaluations are adversarially selected to avoid saturating near perfect performance, so their error rates are not representative of average production traffic.
The clinicians behind the numbers
OpenAI built a Global Physician Network of nearly 300 physicians and psychologists who have practiced in 60 countries to inform its safety research. More than 170 of them — psychiatrists, psychologists, and primary care practitioners — supported this work over the last few months by writing ideal responses for mental health prompts, creating clinically informed analyses of model responses, rating safety across models, and providing high-level guidance.
Psychiatrists and psychologists reviewed more than 1,800 model responses involving serious mental health situations, comparing the new GPT-5 chat model to previous models. They found a 39 to 52 percent decrease in undesired responses versus GPT-4o across all categories — qualitative feedback that echoes the production improvements.
The expert panels are not unanimous, though. OpenAI measured inter-rater agreement at 71 to 77 percent, describing the reliability between expert clinicians as "fair" while acknowledging disagreement in some cases. The company says tracking this variation helps it align model behavior with sound clinical judgment.
What comes next
Going forward, OpenAI is adding emotional reliance and non-suicidal mental health emergencies to its standard baseline safety testing for future model releases, alongside its longstanding suicide and self-harm metrics. The company also collaborated with the Global Physician Network to produce targeted internal evaluations, similar to its HealthBench work, that assess model performance in mental health contexts before release.
OpenAI frames this as unfinished work. "We've made meaningful progress, but there's more to do," the company writes, pledging to keep advancing both its taxonomies and the technical systems used to measure and strengthen model behavior. It also warns that because these tools evolve, future measurements may not be directly comparable to past ones — a note that matters for anyone tracking safety claims across model generations. Further detail appears in an addendum to the GPT-5 system card.
Source: OpenAI News
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
- OpenAI Explains How Its Safety Pipeline Missed GPT-4o Sycophancy
- OpenAI Previews 120-Day Push on ChatGPT Crisis Response and Teen Safety
- OpenAI explains how its own tests missed GPT-4o's sycophancy problem
- OpenAI Ships GPT-5.3 Instant With 26.8% Fewer Hallucinations