OpenAI Gives ChatGPT Cross-Conversation Memory for Crisis Safety
OpenAI says safety updates lifted ChatGPT's safe responses by 50% in suicide/self-harm tests, using cross-conversation "safety summaries" on GPT-5.5 Instant.
Updated
Why it matters
- Safe-response performance improved 50% in suicide/self-harm and 16% in harm-to-others scenarios in long single conversations.
- On GPT-5.5 Instant, multi-conversation safe-response performance improved 52% for harm-to-others and 39% for suicide/self-harm.
- Safety summaries scored 4.93/5 for safety relevance and 4.34/5 for factuality across more than 4,000 evaluations.
- The updates were developed with input from OpenAI's Global Physicians Network over more than two years of expert collaboration.
- OpenAI may extend the same cross-conversation safety methods to biology and cyber safety in the future.
OpenAI says new safety updates have improved ChatGPT's safe-response performance in simulated suicide and self-harm scenarios by 50% in long single conversations — and by 52% in harm-to-others cases across multiple conversations on GPT-5.5 Instant, the current default model in ChatGPT.
The company detailed the changes in a blog post titled "Helping ChatGPT better recognize context in sensitive conversations." The core innovation: ChatGPT can now recognize subtle or evolving cues of risk that emerge over time, both within a single conversation and across separate chats, and use that context to inform safer responses.
The stakes are considerable. OpenAI states that across hundreds of millions of daily interactions, some conversations involve people who are struggling or experiencing distress. The company already provides crisis resources and a "trusted contact" feature that can connect users with someone they trust. The new work aims to catch warning signs that only become visible when messages are read together rather than in isolation.
Why Context Matters in Sensitive Conversations
OpenAI explains that in sensitive conversations, context can matter as much as a single message. A request that appears ordinary or ambiguous on its own may carry a very different meaning when viewed alongside earlier signs of distress or possible harmful intent.
"To respond appropriately, we train ChatGPT to recognize the potential harmful intent from the surrounding context so that it can refuse the request, de-escalate, and guide the user toward support," the company wrote.
The work focuses on acute scenarios: suicide, self-harm, and harm-to-others. OpenAI says these cases are uncommon but critically important to get right. The stated goal is to help ChatGPT connect relevant signals when they matter without overreacting in ordinary conversations.
The approach builds on OpenAI's "safe completion" method, designed to refuse unsafe parts of a user request while responding cautiously where it safely can. The updated model policies and training aim to escalate caution when signals of harm emerge within conversations, while continuing to respond helpfully in benign situations.
What Are Safety Summaries?
The most significant technical change addresses risk that spans separate conversations. One conversation may include subtle signs of potentially harmful intent; another may contain related requests that only trigger concern when combined with the earlier context.
"Without that safety-relevant context, the later conversation – and potentially important warning signs – may appear benign," OpenAI wrote.
To handle this, OpenAI developed safety summaries: short, factual notes about earlier safety-relevant context that may matter in rare, high-risk situations. Their characteristics, per the company:
- Created by a model trained specifically for safety reasoning tasks
- Narrowly scoped and kept only for a limited time
- Used only when relevant to a serious safety concern
- Designed to capture factual safety context, not to serve as general personalization or long-term memory
ChatGPT was also trained to use this context more carefully — for example by de-escalating, refusing to provide details, or redirecting toward safer alternatives.
The design reflects ongoing tension in the industry: retention of context across sessions raises privacy questions even when scoped to safety. OpenAI's framing — narrowly scoped, time-limited, safety-only — reads as a direct answer to that concern, though the company does not specify in the post how long summaries are retained.
How Much Did Performance Improve?
OpenAI reports results from internal evaluations designed to measure performance in challenging cases where risk became clearer over time:
- Long single-conversation scenarios: safe-response performance improved 50% in suicide and self-harm cases and 16% in harm-to-others cases
- Multi-conversation testing on GPT-5.5 Instant: 52% improvement in harm-to-others cases and 39% in suicide and self-harm cases
This means, OpenAI notes, the model was substantially more likely to recognize when earlier parts of a conversation changed the meaning of a later request and respond appropriately.
The company also evaluated the safety summaries themselves. Across more than 4,000 evaluations, the summaries received an average safety relevance score of 4.93 out of 5 and a factuality score of 4.34 out of 5.
OpenAI additionally tested whether adding safety context degraded quality in ordinary conversations. In internal testing, responses remained broadly comparable in everyday chats, with no meaningful user preference between responses with or without safety summaries.
Who Advised the Work?
OpenAI says it developed these systems with input from mental health professionals in its Global Physicians Network, including psychiatrists and psychologists with expertise in forensic psychology, suicide prevention, and self-harm.
These experts helped inform decisions around when safety summaries should be created, how much prior context may be relevant, and how long the model should consider that context when responding.
The updates build on more than two years of collaboration with mental health and safety experts, OpenAI states, spanning model training, evaluations, and monitoring systems. The company's earlier work in this area includes its "strengthening ChatGPT responses in sensitive conversations" initiative and the trusted contact feature.
What Comes Next?
OpenAI frames this as a difficult, long-term challenge: signals of risk can be subtle, spread across messages, or buried within otherwise ordinary conversations.
Today the work covers self-harm and harm-to-others scenarios. In the future, OpenAI says it may explore whether similar methods can help in other high-risk areas such as biology or cyber safety, with careful safeguards in place. The company commits to continuing to strengthen safeguards as its models and understanding evolve — an indication that cross-conversation safety context, now proven in mental-health scenarios, could become a broader safety architecture for future OpenAI models.
Source: OpenAI News
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
191 articles
Related articles
- OpenAI Adds Trusted Contact Alerts to ChatGPT for Self-Harm Safety
- OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
- OpenAI Cuts Unsafe ChatGPT Mental Health Responses by Up to 80%
- OpenAI Details Mental Health Safety Push Across ChatGPT
- OpenAI Ships GPT-5.3 Instant With 26.8% Fewer Hallucinations