Safety & Security

OpenAI Releases MentalHealthBench, an Expert-Built AI Mental Health Benchmark

OpenAI open-sources MentalHealthBench, co-created with 80+ clinicians from 22 countries, to measure AI responses across the full spectrum of mental health conversations.

Introducing MentalHealthBench
Introducing MentalHealthBenchjenschapter3 / Openverse
By Marcus Bennett6 min read

Updated

Why it matters

  • MentalHealthBench was co-created with more than 80 licensed mental health experts from 22 countries, speaking 19 languages and covering nearly 20 subspecialties.
  • Each conversation was reviewed by at least three experts; only criteria agreed by two and not contradicted by a third were kept, with weights from -10 to +10. OpenAI grades responses with GPT‑5.6 Sol.
  • A parallel study with 44 adults across 16 countries found users value practical next steps and tone, while experts emphasize gathering context and interpreting ambiguous situations.

OpenAI has released MentalHealthBench, an open benchmark that measures how AI systems respond across the full spectrum of mental health conversations — from everyday stress to emergencies requiring urgent real-world support. More than 80 licensed psychologists and psychiatrists from 22 countries co-created the benchmark, and OpenAI is releasing it openly so other researchers can inspect the methods, run their own evaluations, and build on the work.

The release addresses a measurable gap in AI evaluation. Most existing assessments of AI in the mental health domain focus on emergency scenarios and measure success using broad, predefined criteria — essentially testing whether a model avoids disallowed responses. That approach leaves open the question of how models perform across the full range of mental health conversations, and how well their responses align with expert guidance for each specific situation. The stakes are considerable: OpenAI says more than one billion people use ChatGPT each week, and many of them turn to the system for conversations about relationships, everyday stress, supporting loved ones, or approaching difficult situations.

How the benchmark works

MentalHealthBench uses privacy-preserving techniques to generate synthetic mental health conversations that OpenAI says accurately reflect real-world usage patterns of AI for mental health support. Some scenarios include relevant background information about the synthetic user — a recent loss in the family, for example — so evaluators can assess whether models use that context to tailor their responses appropriately.

The scenarios cover four user types: adults, teens aged 13–17, caregivers, and clinicians. Conversations span multiple languages and regions and are distributed across the full spectrum of acuity:

  • Non-acute: everyday conversations that may involve some emotional components.
  • High-acuity: conversations indicating more serious mental health concerns or significant distress, but not an immediate emergency.
  • Emergencies: conversations involving signs of a mental health emergency or immediate safety concerns that call for urgent real-world support.

OpenAI notes that the mix of scenarios is designed to stress-test model responses and does not represent how often these topics actually occur in ChatGPT.

Rubrics written by clinicians

The benchmark builds on OpenAI's earlier clinician-informed work on HealthBench and HealthBench Professional. The expert cohort behind MentalHealthBench — more than 80 licensed practitioners speaking 19 languages and representing nearly 20 mental health subspecialties — read each synthetic conversation and produced detailed rubric criteria for evaluating model responses to the final user message.

Each criterion targets a single aspect of a response, such as asking the right question or providing the best possible advice. Criteria carry weights ranging from -10 to +10: positive points reward beneficial behaviors, negative points penalize harmful ones, and criteria with larger values indicate greater clinical importance in the context of a conversation.

The rubric process was deliberately conservative. Each conversation was reviewed by at least three experts, and OpenAI retained only criteria agreed upon by at least two experts and not contradicted by a third. The final rubrics therefore reflect shared clinical judgment about how models should respond in each context.

For scoring, OpenAI uses an automated grader — GPT‑5.6 Sol — to assess model responses against the expert-written criteria. The accompanying paper describes the grading process and evaluation settings in detail.

What the results show

OpenAI evaluated a wide range of recent frontier models on the benchmark, using each provider's newest model as of September 23, 2026. The evaluation measures whether a model's response demonstrates all the ideal behaviors experts identified for each scenario while avoiding less desirable behavior. Error bars in the published results show 95% confidence intervals.

The benchmark's acuity breakdown lets researchers compare model performance on everyday scenarios against urgent emergencies that require real-world support. The teen persona carries an added safeguard consideration: OpenAI explicitly stated through a system message that the user was between 13 and 17, an approach designed to work across model providers, though the company acknowledges it may not capture all safeguards built into individual products. Clinicians with expertise in youth mental health reviewed the teen-persona conversations to assess whether responses fit teenagers' needs.

The benchmark also supports multifaceted measurement. The overall score decomposes into ten dimensions of model behavior defined by the mental health experts, and OpenAI says models with similar overall scores can show different strengths across those dimensions. Fine-grained data of this kind can help researchers pinpoint improvement areas along interpretable behaviors. OpenAI reports one concrete example: a model's ability to seek context appropriately has increased with more advanced models. The company attributes this progress to investments by OpenAI and other model providers in improving how models handle mental health conversations.

OpenAI frames the results with an explicit caveat — ChatGPT is not a substitute for therapy or professional care — and positions the benchmark as a way to measure progress toward models that respond with empathy, promote well-being, and guide people toward real-world support such as localized crisis hotlines or a trusted contact.

Where users and experts disagree

Alongside the benchmark, OpenAI ran a separate analysis comparing expert guidance with what people actually find helpful in AI support. The question was direct: could responses that experts rate highly still feel cold or unhelpful to users?

The study involved 44 adults who had used AI for mental health or emotional support, representing 16 countries and 14 languages. Participants rated model responses to synthetic conversations and wrote their own criteria describing what helpful support should look like. Their review was limited to non-acute conversations to avoid exposing them to potentially distressing high-acuity material. The analysis did not change the benchmark's final scoring criteria, which remain based on expert consensus.

The comparison surfaced a clear divergence. Users valued qualities in AI support that expert guidance emphasized less — particularly practical next steps and tone. Experts, by contrast, placed greater emphasis on gathering relevant context and carefully interpreting ambiguous situations. OpenAI says the comparison gives a fuller picture of what people value in support and how those preferences relate to clinical guidance.

A broader safety push

MentalHealthBench arrives as part of a wider set of OpenAI initiatives at the intersection of AI and mental health. The company is funding new research through grants for AI and mental health, convening experts with the Partnership on AI, and supporting complementary independent efforts such as Transluce's mental health evaluation.

On the product side, OpenAI has strengthened ChatGPT's responses in sensitive conversations, expanded access to crisis resources, and added Trusted Contact to connect people with someone they trust in moments of distress. It has also introduced ChatGPT for Teens, with additional protections for younger users.

"No benchmark captures everything that matters in a personal conversation, but we hope MentalHealthBench helps set a higher standard for how AI supports people," OpenAI writes. The open release invites the research community to examine the methods, identify gaps in existing models, and work toward improving future ones — a signal that OpenAI expects mental health capability, not just safety compliance, to become a standing axis of competition among frontier model providers.

Original: help.openai.com

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

108 articles

Related articles

  1. OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
  2. OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
  3. OpenAI Puts Up to $2 Million Behind AI and Mental Health Research
  4. OpenAI Details Mental Health Safety Push Across ChatGPT
  5. OpenAI Previews 120-Day Push on ChatGPT Crisis Response and Teen Safety

Next article »