OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
OpenAI says GPT-5 updates shaped by 170+ clinicians cut undesired responses in psychosis, self-harm and emotional-reliance conversations by 65-80% in production traffic.

Updated
Why it matters
- OpenAI reports a 65-80% reduction in ChatGPT responses that fall short of desired behavior across mental health-related domains after its latest GPT-5 update.
- More than 170 mental health experts from a Global Physician Network of nearly 300 clinicians across 60 countries supported the research; expert reviews of 1,800+ responses found a 39-52% decrease in undesired responses versus GPT-4o.
- OpenAI is adding emotional reliance and non-suicidal mental health emergencies to its standard baseline safety testing for future model releases.
OpenAI says its latest GPT-5 update has reduced ChatGPT responses that fall short of desired behavior in sensitive mental health conversations by 65% to 80%, according to production traffic measurements the company published in a detailed engineering post. The work drew on more than 170 mental health experts with real-world clinical experience and marks one of the most quantified safety interventions the company has disclosed to date.
The stakes are considerable. ChatGPT serves a growing user base, and a portion of those conversations involve people in acute distress — psychosis, mania, suicidal thinking or unhealthy attachment to the model itself. OpenAI estimates that around 0.15% of users active in a given week have conversations containing explicit indicators of potential suicidal planning or intent, and that 0.07% of weekly active users show possible signs of mental health emergencies related to psychosis or mania. The company frames those figures as its best current estimates, subject to change as measurement methods mature.
The safety work covers three priority domains: mental health concerns such as psychosis and mania; self-harm and suicide; and emotional reliance on AI. Alongside the model changes, OpenAI expanded access to crisis hotlines, re-routed sensitive conversations originating from other models to safer models, and added gentle reminders to take breaks during long sessions.
"We believe ChatGPT can provide a supportive space for people to process what they're feeling, and guide them to reach out to friends, family, or a mental health professional when appropriate," the company wrote.
What changed in the model
The improvements build on OpenAI's Model Spec, the document that defines how its models should behave. OpenAI updated the Spec to make longstanding goals more explicit: the model should support and respect users' real-world relationships, avoid affirming ungrounded beliefs that potentially relate to mental or emotional distress, respond safely and empathetically to potential signs of delusion or mania, and pay closer attention to indirect signals of self-harm or suicide risk.
The engineering process follows five steps: define the harm, measure it, validate definitions with external experts, mitigate through post-training and product changes, then keep measuring and iterating. Central to that process are detailed guides — "taxonomies" — that describe properties of sensitive conversations and what ideal versus undesired model behavior looks like. These taxonomies serve both to train the model and to track its performance before and after deployment.
Because the relevant conversations are extremely rare in production traffic, small measurement differences can materially shift the reported numbers. To compensate, OpenAI runs adversarially difficult "offline evaluations" before deployment, selected for a high likelihood of eliciting undesired responses. The company cautions that these tests are deliberately unsaturated — models don't perform perfectly on them — and their error rates are not representative of average production traffic.
The numbers, domain by domain
For challenging conversations related to mental health issues such as psychosis and mania, OpenAI estimates the latest GPT-5 update cut non-compliant responses by 65% in recent production traffic. On a evaluation set of more than 1,000 challenging mental health-related conversations, automated scoring rates the new GPT-5 model at 92% compliant with desired behaviors, versus 27% for the previous GPT-5 model.
On self-harm and suicide, the company observed an estimated 65% reduction in production traffic in responses that do not fully comply with its taxonomies. Automated evaluations on more than 1,000 challenging conversations score the new model at 91% compliance, up from 77% for the previous GPT-5. OpenAI also reports improved reliability in long conversations — a known failure mode — with its latest models maintaining over 95% reliability in a new set of challenging long conversations built from real-world scenarios selected for their higher likelihood of failure.
The largest gain came in emotional reliance. The company's taxonomy distinguishes healthy engagement from concerning patterns, such as exclusive attachment to the model at the expense of real-world relationships, well-being or obligations. The latest update reduced non-compliant responses by roughly 80% in production traffic. On more than 1,000 challenging conversations indicating emotional reliance, automated evaluations score the new GPT-5 at 97% compliant, compared to 50% for the previous version. OpenAI estimates about 0.15% of weekly active users show potentially heightened emotional attachment to ChatGPT.
Clinicians in the loop
The research rests on OpenAI's Global Physician Network, a pool of nearly 300 physicians and psychologists who have practiced in 60 countries. More than 170 of them — psychiatrists, psychologists and primary care practitioners — supported the work over the last few months by writing ideal responses to mental health-related prompts, producing clinically informed analyses of model outputs, rating the safety of responses across different models, and advising on the overall approach.
In direct comparisons, psychiatrists and psychologists reviewed more than 1,800 model responses involving serious mental health situations. They found the new GPT-5 chat model delivered a 39% to 52% decrease in undesired responses versus GPT-4o across all categories — 39% on challenging mental health conversations (n=677), 52% on self-harm and suicide (n=630) and 42% on emotional reliance (n=507). "In these reviews, clinicians have observed that the latest model responds more appropriately and consistently than earlier versions," OpenAI wrote.
Expert judgment is not uniform. OpenAI measured inter-rater agreement — how often experts reach the same conclusion on whether a response is desirable — and found reliability ranging from 71% to 77%. The company describes that as fair agreement, with visible disagreement in some cases, and says tracking it helps align model behavior with sound clinical judgment.
Why it matters
The disclosure comes as regulators, researchers and courts scrutinize how conversational AI handles vulnerable users, and as competitors face lawsuits over chatbot interactions with minors. By publishing prevalence estimates, evaluation methodology and clinician review data, OpenAI is establishing a template for how frontier labs might be held accountable on mental health safety — while also conceding the fieldwork is still in motion. The company notes that even experts disagree on what the best response looks like in these situations.
Going forward, OpenAI is adding emotional reliance and non-suicidal mental health emergencies to its standard baseline safety testing for future model releases, alongside its longstanding suicide and self-harm metrics. The company has also collaborated with the Global Physician Network on targeted internal evaluations, used to assess models in mental health contexts prior to release — the same playbook it applied to HealthBench.
"We've made meaningful progress, but there's more to do," OpenAI wrote, adding that because taxonomies and measurement systems evolve, future measurements may not be directly comparable to past ones. Further detail appears in an addendum to the GPT-5 system card.
Source: OpenAI News
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles
Related articles
- OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
- OpenAI Explains How Its Safety Pipeline Missed GPT-4o Sycophancy
- OpenAI Releases MentalHealthBench, an Expert-Built AI Mental Health Benchmark
- OpenAI Details Mental Health Safety Push Across ChatGPT
- OpenAI Previews 120-Day Push on ChatGPT Crisis Response and Teen Safety