Models

OpenAI says GPT-5.2 sets new state of the art on FrontierMath

GPT-5.2 Thinking solved 40.3% of FrontierMath problems, a new state of the art. GPT-5.2 Pro hit 93.2% on GPQA Diamond. OpenAI still calls such systems "not independent researchers."

Advancing science and math with GPT-5.2
Advancing science and math with GPT-5.2NASA Goddard Photo and Video / Openverse
By Rebecca Stone3 min read

Updated

Why it matters

  • GPT-5.2 Thinking solved 40.3% of FrontierMath (Tier 1–3) problems, a new state of the art on the expert-level math benchmark.
  • GPT-5.2 Pro scored 93.2% on GPQA Diamond, with GPT-5.2 Thinking at 92.4%; tests ran without tools at maximum reasoning effort.
  • OpenAI explicitly states these systems are "not independent researchers" and that expert judgment, verification, and domain understanding remain essential.

OpenAI says GPT-5.2 Thinking solved 40.3% of problems on FrontierMath (Tier 1–3), an expert-level mathematics evaluation, setting a new state of the art for the benchmark.

The company positions the result as more than a narrow math score. Strong mathematical reasoning, OpenAI argues, lets models follow multi-step logic, keep quantities consistent, and avoid subtle errors that compound in real analyses — from simulations and statistics to forecasting and modeling. "Improvements on benchmarks like FrontierMath reflect not a narrow skill, but stronger general reasoning and abstraction, capabilities that carry directly into scientific workflows such as coding, data analysis, and experimental design," the company wrote in its announcement.

OpenAI frames these capabilities as directly tied to progress toward general intelligence. A system that reasons through abstraction, maintains consistency across long chains of thought, and generalizes across domains exhibits "traits that are foundational to AGI — not task-specific tricks, but broad, transferable reasoning skills that matter across science, engineering, and real-world decision-making."

The benchmark numbers

Beyond FrontierMath, OpenAI reports results on GPQA Diamond, a graduate-level, "Google-proof" multiple-choice benchmark covering physics, chemistry, and biology. GPT-5.2 Pro scored 93.2%, with GPT-5.2 Thinking close behind at 92.4%. The tests ran without tools and with reasoning effort set to maximum.

OpenAI claims GPT-5.2 Pro and GPT-5.2 Thinking are "the world's best models for assisting and accelerating scientists."

The stakes are significant. Labs are competing to prove that frontier models can do useful scientific work — not just chat convincingly — and math benchmarks have become a key proxy for that claim. Research domains with axiomatic foundations, such as mathematics and theoretical computer science, are where OpenAI sees the clearest near-term value: frontier models can help explore proofs, test hypotheses, and identify connections that would otherwise take substantial human effort.

Built on a year of scientist collaboration

The announcement builds on work OpenAI published last month: a paper compiling early case studies across math, physics, biology, computer science, astronomy, and materials science in which GPT-5 contributed to real scientific work. The company says it spent the past year working with scientists across those fields to map where AI helps and where it still falls short. With GPT-5.2, those gains are "becoming more consistent and more reliable," according to OpenAI.

Limits, stated plainly

Notably, OpenAI draws a hard line around what these systems are not. "These systems are not independent researchers," the company states. "Expert judgment, verification, and domain understanding remain essential. Even highly capable models can make mistakes or rely on unstated assumptions."

The company says capable models can still produce detailed, structured arguments that merit careful human study, and that reliable progress depends on workflows keeping "validation, transparency, and collaboration firmly in the loop." Responsibility for correctness, interpretation, and context stays with human researchers, while models like GPT-5.2 serve as tools supporting mathematical reasoning and accelerating early-stage exploration.

The FrontierMath result, viewed as a case study, sketches an emerging division of labor: AI streamlines significant parts of theoretical work while humans keep the central role in scientific judgment. Whether the benchmark gains translate into publishable discoveries at scale remains the open question for the year ahead.

Original: arxiv.org

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Says Two API Settings Tripled GPT-5.6's ARC-AGI-3 Score
  2. OpenAI says GPT-5.6 Sol cut its own serving costs by 20 percent
  3. OpenAI Says 80 to 90 Percent of Its Research Targets GPT 7 and Beyond
  4. OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
  5. OpenAI Says Two API Settings Tripled GPT-5.6's ARC-AGI-3 Score

« Previous articleNext article »