Research

OpenAI says internal model likely solved at least five of ten First Proof research math problems

OpenAI published proof attempts for all ten First Proof research math problems on February 14, 2026. Experts give five a high chance of correctness; one earlier claim has been retracted after community review.

Our First Proof submissions
Our First Proof submissionsDanny Choo / Openverse
By James Calloway6 min read

Updated

Why it matters

  • OpenAI published proof attempts for all ten First Proof problems on February 14, 2026, at 12:00 AM PT, and believes at least five (problems 4, 5, 6, 9, and 10) have a high chance of being correct based on expert feedback.
  • OpenAI initially believed its attempt for problem 2 was correct but now considers it incorrect, citing official First Proof commentary and community analysis.
  • OpenAI researcher James R. Lee said an unreleased model in training solved problems #9 and #10 early and at least three more as training progressed, adding: "It's pretty incredible to watch a model get tangibly smarter day by day."

OpenAI says an internal reasoning model, run with limited human supervision, has produced proof attempts that experts believe have a high chance of being correct for at least five of the ten problems in First Proof, a research-level mathematics challenge designed to test whether AI systems can build complete, checkable arguments in specialized domains.

The company published its proof attempts on Saturday, February 14, 2026, at 12:00 AM PT. Based on feedback from experts, OpenAI believes the attempts for problems 4, 5, 6, 9, and 10 are likely correct, while several others remain under review. The full preprint, containing all ten attempts plus an appendix with prompt patterns and examples simulating the team's manual interactions with the models, is available on OpenAI's website.

One claim has already fallen. OpenAI initially believed its attempt for problem 2 was likely correct. Based on the official First Proof commentary and further community analysis, the company now says it is incorrect. "We're grateful for the engagement and look forward to continued review," the company wrote in its announcement.

Why First Proof matters

First Proof is not a competition-style benchmark. The problems require building end-to-end arguments in specialized domains, and correctness is difficult to establish without expert review — the property that makes the benchmark unusually demanding and unusually informative. The problems were authored by leading experts in their respective fields, and at least a couple of them were open for years before the authors themselves found solutions. The challenge's designers estimate that an academic department with substantial overlap with the subject areas could conceivably solve many of the problems in one week.

That framing sets the bar. A model producing five or more correct proofs in a fast internal sprint would be performing at the level of a concentrated effort by domain specialists, and the fact that one attempt failed under community scrutiny illustrates exactly the kind of failure mode the benchmark is built to expose: reasoning that looks persuasive but does not survive expert examination.

OpenAI argues that this class of challenge is essential for evaluating the next generation of AI models. "We believe novel frontier research is perhaps the most important way to evaluate capabilities of next generation AI models," the company wrote. Benchmarks are useful, but they can miss some of the hardest parts of research: sustaining long chains of reasoning, choosing the right abstractions, handling ambiguity in problem statements, and producing arguments that survive expert scrutiny. Frontier challenges like First Proof, the company says, stress-test those capabilities in settings where correctness is nontrivial to verify and the failure modes are informative.

"Pretty incredible to watch a model get tangibly smarter day by day"

The most striking detail in OpenAI's account is the trajectory of the model during training. James R. Lee, a researcher working on reasoning at OpenAI, described running the benchmark on the model as it trained:

"We're currently training a new model for which a primary focus is increasing the level of rigor in its thinking, with the goal that the model can think continuously for many hours and remain highly confident in its conclusions. When the First Proof problems were announced, it seemed like the perfect testbed, so over the weekend I tried it out. Already it was able to solve two of the problems (#9 and #10). As it trained, it became increasingly capable, eventually solving–in our estimation–at least three more. We were particularly pleased when it solved #6 and then, two days later, #4, as those problems were from fields familiar to many of us. It's pretty incredible to watch a model get tangibly smarter day by day."

The quote is notable for two reasons. First, it confirms the model in question is an unreleased system still in training, with rigor in long-form reasoning as a primary objective. Second, it documents measurable capability gains on fixed, expert-authored problems across a window of days — problems 9 and 10 fell early, problem 6 and then problem 4 came two days apart as training progressed.

The process was messier than a controlled evaluation

OpenAI is candid that the results do not come from a clean experimental setup. The model ran with limited human supervision, and the team's interventions went beyond simple prompting. When prompting versions of the model along training, researchers sometimes suggested retrying strategies that appeared fruitful in earlier attempts. For some attempts, they asked the model to expand or clarify parts of a proof after receiving expert feedback, to make the reasoning easier to verify. The team also facilitated a back-and-forth between the internal model and ChatGPT for verification, formatting, and style. For some problems, OpenAI presents the best of a few attempts, selected by human judgment.

"This was a fast sprint, and our process was not as clean as we would like in a properly controlled evaluation," the company acknowledged. OpenAI says it looks forward to discussions with the First Proof organizers about a more rigorous experiment and evaluation framework for future iterations.

That admission matters for how the results should be read. The five-attempt count reflects OpenAI's own assessment informed by expert feedback, not a formal adjudication, and the retraction of the problem 2 attempt shows how fluid the picture remains. Independent verification by the problem authors and the broader mathematics community is still the decisive test.

A line of escalating results in research-grade math and science

The First Proof results extend a progression OpenAI has documented over the past year. In July 2025, the company reported gold medal-level performance on the International Mathematical Olympiad with a general-purpose reasoning model, scoring 35 out of 42 points. In November 2025, it published "Early experiments in accelerating science with GPT-5," a set of case studies in which GPT-5 helped researchers make concrete progress across math, physics, biology, and other fields — along with the limitations the company observed. Most recently, OpenAI reported a physics collaboration in which GPT-5.2 proposed a candidate expression for a gluon-amplitude formula that was then formally proved by an internal model and verified by the authors.

The through-line is a shift in evaluation targets: from competition problems with known answers, to open research questions where the model must construct novel arguments that experts can check. First Proof sits at the hard end of that spectrum, and the stakes for the field are straightforward. If frontier models can reliably produce correct research-level proofs, they become credible collaborators for working mathematicians rather than benchmark performers. If they cannot, failures like the retracted problem 2 attempt will keep defining the gap.

OpenAI says it wants deeper engagement with the community on how to evaluate research-grade reasoning, including expert feedback on the published attempts, and the company states it is excited to make these new capabilities available in future public models. The next milestone to watch is whether the five believed-correct attempts survive formal review — and whether the still-training model, which Lee watched solving additional problems over a single weekend, appears in a release with the rigor-focused capabilities the company describes.

Original: 1stproof.org

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

121 articles

Related articles

  1. OpenAI Says Internal Model Has Solved Over 100 Open Math Problems
  2. OpenAI Recruits Elite Mathematicians After Research Release Stumbles
  3. OpenAI's Math Advisory Group Off to Another Rocky Start
  4. OpenAI says GPT-5.2 sets new state of the art on FrontierMath
  5. OpenAI Launches FrontierScience Benchmark for AI Research Skills

« Previous articleNext article »