OpenAI's o3 Helps Diagnose 18 Rare Disease Cases That Stumped Specialists
OpenAI o3 Deep Research helped clinicians diagnose 18 of 376 unsolved rare disease cases, a 4.8% additional yield, in a study published June 18, 2026 in NEJM AI.

Updated
Why it matters
- Researchers used OpenAI o3 Deep Research on 376 unsolved rare disease cases; physicians confirmed 18 diagnoses, an additional yield of 4.8% after prior specialist review.
- The study, published June 18, 2026 in NEJM AI, involved Boston Children's Hospital's Manton Center, Harvard University, and OpenAI.
- The model never made a diagnosis: every finding passed through ACMG/AMP expert review, CLIA-certified laboratory confirmation, and clinical return to families.
- Next-stage work will be led by the Manton Center under an OpenAI Foundation grant to build a platform-agnostic genetics AI copilot.
An AI reasoning model helped physicians diagnose 18 children and young adults whose rare genetic diseases had survived years of specialist review, according to a study published June 18, 2026 in NEJM AI.
Researchers from Boston Children's Hospital's Manton Center for Orphan Disease Research, Harvard University, and OpenAI used the OpenAI o3 Deep Research reasoning model to reanalyze de-identified clinical and genomic data from 376 previously analyzed but unsolved cases. After the model surfaced candidate explanations, experts reviewed them, ordered additional testing where appropriate, and confirmed findings through standard clinical processes. The result: an additional diagnostic yield of 4.8% on cases that earlier analysis by specialists had left unresolved.
The stakes are large. Roughly half of people with rare diseases never receive a clear genetic diagnosis, even after genomic sequencing, extensive testing, and specialist review. Their medical data may hold clues, but finding them can require sifting through thousands to millions of possible genetic variants, fragmented clinical records, and a scientific literature that changes monthly. As new gene-disease relationships, case reports, and classification evidence accumulate, unsolved cases can become newly interpretable — but someone, or something, has to go back and look.
That is the maintenance problem this study attacks. A patient's genome stays fixed; the evidence around it does not. Researchers link new genes to disease, labs reclassify old variants, and databases accumulate observations. Each update can make an old inconclusive case worth revisiting, and many institutions now carry a growing backlog of genomes that need resynchronization with a moving knowledge base. Periodic expert-led reanalysis works, but it is expensive and does not scale. The study's central question is whether a general-purpose reasoning model can widen that search cheaply while leaving every clinical decision in human hands.
An explanation-first reasoning layer
The team, led in the Manton Center by Dr. Catherine Brownstein and Alan Beggs, director of the Manton Center for Orphan Disease Research, designed the workflow so the model acted as an explanation-first reasoning layer on top of existing genomic pipelines. The model did not return a bare ranked gene. It was asked to connect clinical features, inheritance pattern, variant evidence, and scientific literature into a justification a human reviewer could interrogate.
For each case, researchers assembled a de-identified packet: standardized Human Phenotype Ontology terms describing the patient's clinical presentation, occasional clinician notes and descriptive diagnoses, metadata such as age and gender, and a filtered variant table capturing each variant's rarity, predicted protein effect, ClinVar classification, and signal quality across family members. Most cases included data from the child and both biological parents.
The guardrails were strict. Researchers reviewed the model's outputs using the same ACMG/AMP framework clinical labs use to classify variants. At least two team members reviewed each candidate. Disagreements were resolved by consensus. A model output was never treated as a diagnosis. A finding counted only after qualified experts reviewed the evidence, the variant was classified pathogenic or likely pathogenic, a CLIA-certified laboratory confirmed it, and the clinical team returned the result to the family.
Calibrating the model on solved cases first
Before touching unsolved cases, the team tuned the workflow on cases with established diagnoses. In duplicate runs, it recovered the correct gene and variant in 48 of 51 cases spanning a variety of rare conditions. In a set of 57 neuromuscular cases, it returned the correct diagnosis in duplicate runs for 45. In a 15-case long-read genome set, it named the correct gene in every case and both disease-causing alleles in 12.
The model's self-reported confidence scores tracked with correct answers in these previously solved cases: the mean minimum score was 85.6 for consistently correct calls and 42.1 for incorrect or unknown calls. The scores were not calibrated probabilities, and the team did not use them as a substitute for evidence or clinical adjudication. They did help expert reviewers prioritize the most promising candidates.
Results across four difficult cohorts
The team then applied the workflow to four groups of previously unsolved cases: children with neurodevelopmental conditions, people with rare neuromuscular disease, children and adolescents with early psychosis, and cases of sudden unexpected death in pediatrics. These were not fresh cases awaiting first review. Many had already been examined by multiple commercial or institutional pipelines and discussed by multidisciplinary teams.
| Cohort | Cases | Diagnoses surfaced | Yield |
|---|---|---|---|
| Neurodevelopmental | 100 | 10 | 10.0% |
| Neuromuscular disease | 61 | 4 | 6.6% |
| Sudden unexpected death in pediatrics | 200 | 2 | 1.0% |
| Early psychosis | 15 | 2 | 13.3% |
| Total | 376 | 18 | 4.8% |
The early psychosis cohort was small, so its percentage carries a wide confidence interval, and yield reflects how likely each cohort was to have a single-gene explanation at all.
The 4.8% rate is modest but meaningful in this population. Previous expert reviews had not resolved these cases. Similar reanalysis studies report single-digit gains in heavily reviewed cases; higher yields usually come from studies containing new cases or well-known disorders awaiting genetic confirmation.
Notably, 7 of the 18 diagnoses were rediscoveries — diagnoses established outside the local research workflow but absent from the record the team reviewed. In several of those, the variants were already listed as pathogenic or likely pathogenic in public databases. That detail highlights the operational challenge of synthesizing information across fragmented data sources, which is itself part of why these cases went unsolved.
Reasoning beyond the input data
In one early-psychosis case, the model inferred a structural event in the genome that was not listed in the input data. It connected a run of low-quality calls on chromosome 22 with the child's cardiac, immune, neurodevelopmental, and psychiatric features, then hypothesized a 22q11.2 deletion associated with DiGeorge syndrome. Follow-up genome sequencing confirmed it.
The model also flexed beyond single-gene answers. Although the prompt asked for one monogenic cause, it sometimes surfaced two genes that better explained a complex presentation. Variants in LAMA2 and FOXP1 together accounted for muscle and neurodevelopmental features in one case; another had a previously unrecognized digenic explanation involving TTN and SRPK3.
It produced at least one genuinely novel scientific lead. In a neurodevelopmental case, the model flagged an 11-amino-acid deletion in S1PR1 in a person with vitiligo. S1PR1 encodes a cell-surface receptor involved in signaling, immune-cell movement, and tissue biology. The model integrated evidence suggesting the deletion could alter receptor structure and signaling in ways that reduce pigment production while helping immune cells persist in the skin. The proposed S1PR1-vitiligo relationship requires experimental validation, but it shows how an AI system can translate scattered findings from structural biology, immunology, and clinical genetics into a concrete, testable hypothesis. The team also saw possible phenotype expansion in the neuromuscular cohort, where damaging variants in HSPB8 and CDK13 did not perfectly match the genes' best-known disorders.
What the study does not claim
The researchers are explicit about the boundaries. The study is retrospective, the cohorts were heterogeneous, and reviewers were not blinded to model confidence. The team did not measure time saved, cost, clinician effort, false-positive workload, or changes in care. It did not systematically evaluate structural variants, repeat expansions, deep-intronic changes, or mosaicism.
Large language models can misread context or produce plausible explanations that fail under inspection, which is why every result passed through human adjudication and clinical confirmation. The study states plainly that it is not evidence that patients, clinicians, or customers should use OpenAI models to diagnose disease or make medical decisions, and it does not describe or endorse any intended customer use of OpenAI o3 Deep Research, ChatGPT, or any other OpenAI product for diagnosis. The model did not diagnose any participant. Physicians and qualified clinical experts made every diagnosis through established review, testing, and confirmation processes.
The work used de-identified information, with no protected health information transmitted outside approved environments. The researchers note that broader deployment will demand the same attention to privacy, security, auditability, and local regulation that applies to all medical care — and that model access does not replace sequencing infrastructure, genetic counseling, confirmatory testing, or specialist judgment.
What comes next
The next stage has funding. OpenAI helped support the initial study, but the Manton Center will lead the follow-on work through a grant from the OpenAI Foundation. The grant will support the Center's effort to build a platform-agnostic, low-cost genetics AI copilot that helps clinical teams analyze rare disease cases more quickly and consistently.
The researchers call for prospective, multi-center studies comparing LLM-assisted reanalysis with standard practice on diagnostic yield, time to a candidate, clinician effort, false-positive burden, cost, and effects on care. Versioned prompts, reference checks, audit logs, and calibrated uncertainty will matter for reproducibility and safety. Newer general-purpose models can search and synthesize more scientific material, and purpose-built systems such as GPT-Rosalind target deeper life-sciences work including variant effects on protein structure and function — capabilities not tested here.
The promise the study's authors point to is narrow and concrete: not that AI replaces a doctor's diagnosis, but that carefully evaluated research tools may help specialists identify evidence worth investigating. For thousands of families with unanswered questions, today's dead ends do not have to stay dead forever.
Source: OpenAI News
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
119 articles
Related articles
- Boston Children's says AI has unlocked 40-plus rare disease diagnoses
- Google DeepMind Unveils AI Co-Clinician Research Initiative
- OpenAI Launches GPT-Rosalind, a Reasoning Model for Life Sciences
- DeepMind's AlphaGenome Atlas Maps All 9 Billion Human DNA Variants
- OpenAI launches misalignment disclosure framework, publishes six reports