OpenAI Says GPT-5.5 Instant Now Rival Doctors on Health Questions
GPT-5.5 Instant matches frontier Thinking models on health benchmarks and outscored physician-written responses in a blind review, per OpenAI. 230M use ChatGPT weekly for health questions.

Updated
Why it matters
- GPT-5.5 Instant (released May 2026, free tier) matches frontier Thinking models on OpenAI's hardest health evaluations, including HealthBench Professional.
- In a blind review of 3,500 responses, physicians rated GPT-5.5 Instant higher than physician-written responses and older models across accuracy, communication, and helpfulness.
- Flagged health factuality issues in production traffic (billions of messages weekly) fell 71% in two months; over 260 physicians across 60 countries have reviewed 700,000+ model responses.
OpenAI's GPT-5.5 Instant, a model available to all free ChatGPT users, now performs on the company's most challenging health evaluations at a level comparable to its frontier Thinking models — and, in a blind physician comparison, its responses were rated higher than those written by doctors with unlimited time and internet access.
The company disclosed the results in a blog post titled "Improving health intelligence in ChatGPT," published alongside new details about the physician network behind its health evaluation program. The stakes are considerable: OpenAI says more than 230 million people use ChatGPT every week for health and wellness questions, ranging from interpreting lab results and preparing for appointments to navigating insurance and building healthier habits.
GPT-5.5 Instant, released in May 2026, replaced GPT-5.3 Instant (released March 2026) as the free-tier model. OpenAI reports substantial improvements over its predecessor in recognizing when urgent care may be needed, asking for relevant context, explaining uncertainty, and making complex information easier to understand. The company used API pricing to calculate costs for its 5.4 Thinking and 5.5 Thinking frontier models, which served as the comparison point for the aggregate health benchmark results.
How OpenAI measured it
The evaluation relies on two health-specific benchmarks: HealthBench and HealthBench Professional. Both use realistic health conversations and physician-written rubrics to score responses on accuracy, safety, communication, context awareness, completeness, and appropriate escalation — whether the model tells a user to seek care when they should. On an aggregate of health evaluations including HealthBench Professional, GPT-5.5 Instant reached performance similar to OpenAI's latest frontier models, a substantial jump from GPT-5.3 Instant.
The second comparison is more provocative. OpenAI asked physicians to write responses for representative health conversations, giving them unlimited time and internet access but no AI tools. A separate panel of physicians then compared those physician-written responses with model responses over time, scoring them on accuracy, communication, completeness, instruction following, and health decision helpfulness across 3,500 reviewed responses. GPT-5.5 Instant responses were rated higher than both the physician-written responses and older model outputs across the criteria.
The physicians also tracked failure modes. GPT-5.5 Instant had fewer instances of not tailoring responses to local healthcare context, missing red flags or failing to refer the user to care, and not seeking additional context when needed — fewer than both older models and the physicians themselves.
A third data point comes from production traffic. OpenAI says it uses privacy-preserving monitors on live health conversations — billions of messages a week — to track possible factuality issues. According to the company, the rate of health responses with at least one flagged factuality issue has fallen by 71% over the last two months.
The physician network behind the numbers
The results rest on an evaluation apparatus OpenAI has built over time. The company works with a global network of more than 260 physicians across 60 countries, 49 languages, and 26 medical specialties. Their brief is to define what "good" looks like in real-world health situations: they review example model responses, describe ideal behavior, and identify failure modes.
The scale of the review effort is large. Physicians have reviewed more than 700,000 example model responses reflecting how patients and clinicians actually use ChatGPT, and OpenAI says a physician reviews a new response every few minutes. Their feedback becomes rubrics and evaluation criteria that researchers use to measure whether responses are accurate, safe, clear, complete, appropriately cautious, and useful — giving the company a defined way to track where models improve and where they still fall short.
Physicians specifically look for responses that miss important context, sound too confident, are unclear about next steps, or fail to directly encourage someone to seek medical care.
Why it matters
Health is emerging as one of the highest-stakes arenas for consumer AI. With 230 million weekly health-related interactions, ChatGPT is already functioning as de facto health infrastructure for a large share of its user base, and small differences in escalation judgment — whether a model recognizes a red flag — carry real consequences. The claim that a free, instant-tier model now matches frontier models on health benchmarks narrows the gap between paying and free users on precisely the domain where accuracy matters most.
The physician-comparison result will also draw scrutiny. Blind comparisons of model versus physician responses have become a contested genre in AI evaluation, and OpenAI's methodology — physician raters scoring responses on communication and helpfulness alongside accuracy — leaves open questions about how the results would hold up under independent review. The company notes that "improving human health will be one of the most personal, tangible impacts of AGI."
The health work feeds OpenAI's broader clinical push, including ChatGPT for Clinicians and OpenAI for Healthcare, which support medical professionals with documentation, research, and care delivery. OpenAI says its goal is to keep making ChatGPT more accurate, more useful, and more impactful in health moments — and to keep bringing that progress to more people, including the free tier where GPT-5.5 Instant already ships.
Original: arxiv.org
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
139 articles
Related articles
- OpenAI ships GPT-5.5 Instant with 52.5% fewer hallucinations
- OpenAI Updates GPT-5.6 Sol and Gives Free Users Unlimited Chats
- OpenAI Ships GPT-5.3 Instant With 26.8% Fewer Hallucinations
- OpenAI Launches ChatGPT Go Worldwide With GPT-5.2 Instant Access
- OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price