Research

OpenAI Claims GPT-5 Helped Close an Erdős Problem and Speed Immunology Breakthroughs

OpenAI's new paper, co-authored with university and national lab scientists, details cases where GPT-5 completed proofs, decoded immune-cell data in minutes, and recovered black hole symmetries.

Early experiments in accelerating science with GPT-5
Early experiments in accelerating science with GPT-5Daniel Mennerich / Openverse
By Rebecca Stone6 min read

Updated

Why it matters

  • OpenAI released 'Early science acceleration experiments with GPT-5' on November 20, 2025, co-authored with researchers from Vanderbilt, UC Berkeley, Columbia, Oxford, Cambridge, Lawrence Livermore National Laboratory, and The Jackson Laboratory.
  • GPT-5 contributed the key insight that let Mehtaab Sawhney and Mark Sellke complete a proof of Erdős Problem 848, and independently confirmed Derya Unutmaz's immunology hypothesis from unpublished data within minutes.
  • OpenAI concedes the case studies are curated, not systematic, and that GPT-5 can hallucinate citations, mechanisms, and proofs, and missed prior published work behind one of its own results.

OpenAI has published a paper claiming that GPT-5 helped researchers solve a decades-old problem posed by Paul Erdős, identify the mechanism behind a puzzling immune-cell shift in minutes, and reconstruct hidden symmetries of rotating black holes — results the company says mark the point where frontier models move beyond summarizing existing knowledge and start contributing to new science.

The paper, "Early science acceleration experiments with GPT-5," went live on November 20, 2025. OpenAI co-authored it with researchers at Vanderbilt, UC Berkeley, Columbia, Oxford, Cambridge, Lawrence Livermore National Laboratory, and The Jackson Laboratory. It compiles case studies spanning mathematics, physics, biology, computer science, astronomy, and materials science, and it documents the model's limitations alongside its wins.

The stakes are straightforward. OpenAI cites a recent survey in which 60 percent of people in the U.S. said scientific and medical breakthroughs reach them too slowly, 73 percent said better ways to accelerate discovery are needed, and 69 percent identified scientific leadership as a top national priority. If AI can compress the time between idea and tested result, the company argues, the benefits compound across health, energy, and security research.

Three headline results

The most concrete claims involve named researchers and checkable outcomes.

In biology, Derya Unutmaz, M.D., and his lab had spent months trying to explain a lasting shift in human CD4+ T cells toward a proinflammatory Th17-like state after transient treatment with 2-deoxyglucose (2DG), a compound that interferes with glucose metabolism. Years later, Unutmaz gave GPT-5 Pro an unpublished figure of flow cytometry scatterplots and asked what might explain the data and what experiments to run next. In about a dozen minutes of back-and-forth, the model proposed that disrupted N-linked glycosylation during priming was the driver, predicted that memory rather than naïve T cells were responsible, and suggested follow-up experiments — including a mannose rescue experiment. The lab had already run that experiment, and OpenAI reports the results "matched the model predictions exactly." GPT-5 Pro also predicted that transient 2DG exposure during CAR-T generation would enhance killing efficiency against target cancer cell lines, a prediction the company says matched the lab's unpublished data.

In mathematics, Mehtaab Sawhney and Mark Sellke attacked Erdős Problem 848: finding the largest set of positive integers where, for any two numbers, their product plus one is always divisible by a perfect-square prime factor. The pair had explored the structure but were stuck on the final step. GPT-5 suggested a clearer way to show that a single out-of-place number forces contradictions across almost all other numbers in the set. That idea completed the proof, confirming Erdős's original guess.

In optimization, Microsoft researcher-turned-academic Sébastien Bubeck gave GPT-5 a weaker version of a recent convex optimization theorem by Guy Barzilai, Ohad Shamir, and Moslem Zamani about when gradient descent values form a convex curve over time. The model proposed a sharper step-size bound and a cleaner proof, which Bubeck verified by hand. With more thinking time, an internal run of the model derived the optimal bound from scratch.

Symmetries, literature search, and a cautionary tale

Several case studies illustrate narrower but still notable capabilities.

Physicist Alex Lupsasca had recently shown that the Kerr black hole wave equation carries a hidden symmetry structure forming an SL(2,ℝ) algebra. Asked directly about the full Kerr problem, GPT-5 Pro initially failed and reported no interesting symmetries. After Lupsasca supplied a simpler "warm-up" version in flat space, the model returned to the Kerr case and, after roughly 18 minutes of internal reasoning, produced the full set of symmetry generators closing into SL(2,ℝ), matching the human result.

The paper also flags attribution as an open problem. In one project on error-correcting codes designed to exclude clique-like codewords, GPT-5 reframed the question using quadratic equations over a finite field and invoked the Chevalley–Warning theorem, showing that only about half as many parity-check constraints were needed as previously thought. The catch: the same bound and essentially the same proof had appeared years earlier in a short paper. GPT-5 reproduced the argument without citing its source, and only identified the prior work when asked again in a fresh session. "Models can generate correct and elegant reasoning, but they may not reliably attribute where those ideas originally came from," OpenAI writes.

Fields Medalist Tim Gowers ran his own experiments, treating GPT-5 as a "research partner" on hard combinatorics questions. In multiple cases the model quickly spotted flaws or missing cases in candidate constructions and proposed simpler alternatives or counterexamples; in others it stalled. Gowers's conclusion, as reported by OpenAI: the model is already useful as a very fast, very knowledgeable critic, though it does not yet meet his bar for full co-authorship.

Other case studies include GPT-5 tightening a lower bound for the convex body chasing problem in online algorithms with Christian Coester; producing short, self-contained proofs of two graph-theory inequalities in trees — including one that had been only conjectured — which Bubeck, Sellke, and Yin checked and adopted; helping cosmologist Robert Scherrer catch algebraic slips and translate between dark energy parameterizations; supporting fusion and plasma physics simulations at building and interpreting a reaction–diffusion model of thermonuclear burn propagation; and helping Nikita Zhivotovskiy connect a new convex geometry theorem to density estimation, learning theory, and multi-objective optimization, surfacing references in languages he had not encountered.

Sawhney and Sellke also used GPT-5 as a literature-search assistant over the public Erdős problem database, where some problems remain listed as open despite existing solutions in obscure journals. The model found full solutions for several problems still marked open, identified substantial partial results for others, and flagged a misprint in one problem statement.

The honest section: limitations

OpenAI is explicit that the case studies are "curated illustrations" rather than a systematic sample, and that they do not capture the full range of failure modes. The company lists what can go wrong: GPT-5 sometimes hallucinates citations, mechanisms, or proofs that appear plausible; it is sensitive to scaffolding and warm-up problems, as the black hole case showed; it misses domain-specific subtleties; and it can follow unproductive lines of reasoning if not corrected. Expert oversight remains essential throughout.

The paper's framing also draws a line between specialized scientific tools — simulation engines, protein databases, computer algebra systems — and general foundation models. OpenAI's stated approach uses the former where they exist and builds general reasoning models where they do not, with the two paths reinforcing each other.

Why it matters

The release functions as both a research report and a positioning statement for OpenAI for Science, the initiative behind the work, whose mission is to accelerate discovery by pairing frontier models with tools, workflows, and collaborations across academia, industry, and national labs. The paper does not claim autonomy. It claims leverage: GPT-5, used by experts, expands the surface area of exploration and shortens parts of the research workflow.

The forward-looking claim is the sharpest one in the document. OpenAI notes a trajectory in which these systems improve with more time and compute: if GPT-5 can meaningfully assist with some research questions in 20 minutes, the company expects deeper results when models can spend hours or days reasoning about a problem. Combined with what it calls world-class scientists, OpenAI sees the possibility of a step-change in scientific productivity over time — a claim the wider research community will now test against its own replications.

Original: openaifoundation.org

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Says 80 to 90 Percent of Its Research Targets GPT 7 and Beyond
  2. OpenAI says GPT-5.2 sets new state of the art on FrontierMath
  3. OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
  4. OpenAI Launches GPT-Rosalind, a Reasoning Model for Life Sciences
  5. OpenAI Signs Memorandum of Understanding With U.S. DOE

« Previous articleNext article »