OpenAI Updates GPT-5 System Card With New Mental Health Evals
OpenAI's GPT-5 system card addendum adds safety metrics for GPT-5.1 Instant and Thinking, introducing first-time evaluations for mental health and emotional reliance.

Updated
Why it matters
- OpenAI published an addendum to its GPT-5 system card covering GPT-5.1 Instant and GPT-5.1 Thinking.
- The addendum provides updated safety metrics for both models.
- It introduces new evaluations for mental health and emotional reliance — a new category for the GPT-5 family.
OpenAI has published an addendum to its GPT-5 system card, extending the company's published safety record to cover two newer models: GPT-5.1 Instant and GPT-5.1 Thinking.
The document's core purpose is refreshingly narrow. It supplies updated safety metrics for both GPT-5.1 variants and introduces a new category of evaluation that OpenAI has not previously reported against its GPT-5 models: mental health and emotional reliance.
That addition matters. System cards are OpenAI's primary mechanism for disclosing how frontier models perform on safety evaluations before and after release, covering areas such as harmful content generation, deception, and autonomous capabilities. By appending an addendum rather than issuing an entirely new card, OpenAI is treating GPT-5.1 Instant and GPT-5.1 Thinking as iterations within the GPT-5 family — models close enough to their predecessor that incremental metric updates suffice, but distinct enough that the safety record needs refreshing.
The more consequential signal is what the new evaluations target. Mental health and emotional reliance assessments probe a failure mode that regulators, researchers, and child-safety advocates have flagged repeatedly as AI companionship products reach hundreds of millions of users: models that respond in ways which could worsen psychological distress or foster unhealthy dependency. The addendum brings these concerns formally into OpenAI's published evaluation framework for the GPT-5 family.
The split between the two models is also worth noting. GPT-5.1 Instant and GPT-5.1 Thinking represent different points on the speed-versus-reasoning trade-off — one optimized for fast responses, the other for extended deliberation. Publishing separate safety metrics for each signals that OpenAI considers the distinction material to safety outcomes, not merely to latency or pricing. A reasoning model that deliberates longer can produce different behaviors on sensitive topics than a model tuned for immediacy, and the addendum's structure reflects that.
For enterprises and developers building on OpenAI's API, system card addenda serve as due-diligence artifacts. They are the documents procurement teams, compliance officers, and auditors cite when assessing whether a model change introduces new risk. An addendum covering mental health and emotional reliance gives those stakeholders a named, dated reference point for exactly the questions they are most often asked about consumer-facing deployments in health-adjacent and companion contexts.
The format also sets a precedent. If OpenAI continues to publish addenda for point releases rather than reserving full system cards for major version jumps, the industry gets a faster cadence of safety disclosure — one matched to the actual pace of model shipping, which has accelerated well beyond the annual-or-slower release cycles that full system cards were originally designed around.
Source: OpenAI News
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles