OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
GPT-5 hits 94.6% on AIME 2025 and 74.9% on SWE-bench Verified, cuts hallucinations sharply, and replaces five prior models as ChatGPT's default.

Updated
Why it matters
- GPT-5 is now ChatGPT's default model, replacing GPT-4o, o3, o4-mini, GPT-4.1, and GPT-4.5.
- State-of-the-art results: 94.6% on AIME 2025, 74.9% on SWE-bench Verified, 88.4% on GPQA (GPT-5 pro).
- GPT-5 thinking produces roughly six times fewer hallucinations than o3 on public factuality benchmarks.
OpenAI has released GPT-5, a unified AI system that posts state-of-the-art benchmark scores of 94.6% on AIME 2025 math without tools, 74.9% on SWE-bench Verified coding, and 88.4% on GPQA science questions via the GPT-5 pro variant. The model is now the default in ChatGPT, replacing GPT-4o, OpenAI o3, OpenAI o4-mini, GPT-4.1, and GPT-5's predecessor GPT-4.5 for signed-in users.
The release matters beyond the numbers. OpenAI is consolidating what had become a confusing lineup of models into a single system that decides on its own when to answer quickly and when to reason at length — a bet that automatic routing, not model pickers, is how consumers and enterprises will interact with AI going forward.
One system, three layers
GPT-5 consists of a smart, efficient model that handles most questions, a deeper reasoning model the company calls GPT-5 thinking for harder problems, and a real-time router that chooses between them based on conversation type, complexity, tool needs, and explicit user intent — typing "think hard about this" in a prompt forces reasoning. The router trains continuously on real signals, including when users switch models, response preference rates, and measured correctness. When users hit usage limits, a mini version of each model takes over. OpenAI says it plans to merge these capabilities into a single model in the near future.
Efficiency claims
GPT-5 extracts more performance from less compute at inference time. According to OpenAI's evaluations, GPT-5 with thinking outperforms o3 while using 50-80% fewer output tokens across capabilities including visual reasoning, agentic coding, and graduate-level scientific problem solving. The model was trained on Microsoft Azure AI supercomputers.
Hallucination and honesty improvements
OpenAI claims substantial reliability gains. With web search enabled on anonymized prompts representative of ChatGPT production traffic, GPT-5's responses are roughly 45% less likely to contain a factual error than GPT-4o. When thinking, they are about 80% less likely to contain a factual error than o3. On two public factuality benchmarks, LongFact and FActScore, GPT-5 thinking produced about six times fewer hallucinations than o3.
The honesty results are more striking. To test whether reasoning models lie about completing impossible tasks, OpenAI removed all images from prompts in the multimodal benchmark CharXiv. OpenAI o3 gave confident answers about non-existent images 86.7% of the time. GPT-5 did so just 9% of the time. On a large set of conversations representative of real production traffic, OpenAI reduced deception rates from 4.8% for o3 to 2.1% for GPT-5 reasoning responses. The company acknowledges more work remains.
A new approach to safety training
GPT-5 also marks a shift in how OpenAI handles safety. Instead of relying primarily on refusal-based training — comply or refuse — the company introduced what it calls "safe completions," which teaches the model to give the most helpful answer possible within safety boundaries, sometimes partially answering or answering only at a high level. When the model must refuse, it is trained to explain why and offer safe alternatives. OpenAI says the approach enables better handling of dual-use domains such as virology, stronger robustness to ambiguous intent, and fewer unnecessary overrefusals.
The stakes are highest in biological risk. OpenAI classified GPT-5 thinking as High capability in the biological and chemical domain and completed 5,000 hours of red-teaming with partners including CAISI and the UK AISI under its Preparedness Framework. The company says it has no definitive evidence the model could meaningfully help a novice create severe biological harm — its defined threshold for High capability — but is activating safeguards precautionarily, including comprehensive threat modeling, safe completions training, always-on classifiers, reasoning monitors, and enforcement pipelines.
Coding, writing, and health
OpenAI calls GPT-5 its strongest coding model to date, with particular improvements in complex front-end generation and debugging larger repositories. Early testers noted better design choices around spacing, typography, and white space, and the company demoed a single-prompt game, "Jumping Ball Runner," built in one HTML file with parallax scrolling and high-score tracking.
On health, GPT-5 scores 46.2% on HealthBench Hard, significantly higher than any previous model on the evaluation OpenAI published earlier this year. OpenAI frames the model as a partner that flags potential concerns and adapts answers to a user's context, knowledge level, and geography — not a replacement for medical professionals.
On economically important knowledge work, OpenAI reports that GPT-5 using reasoning is comparable to or better than experts in roughly half the cases across tasks spanning over 40 occupations, including law, logistics, sales, and engineering, while outperforming o3 and ChatGPT Agent.
Sycophancy cut by more than half
After a GPT-4o update earlier this year unintentionally made the model overly agreeable — a change OpenAI quickly rolled back — the company built new evaluations to measure sycophancy. In targeted evaluations using prompts designed to elicit sycophantic responses, GPT-5 cut sycophantic replies from 14.5% to less than 6%. OpenAI says the model is "less effusively agreeable," uses fewer unnecessary emojis, and should feel less like "talking to AI" and more like chatting with a helpful friend with PhD-level intelligence.
The improved steerability also enables a research preview of four preset personalities for all ChatGPT users: Cynic, Robot, Listener, and Nerd. The personalities are opt-in, adjustable in settings, initially text-only, and meet or exceed OpenAI's internal bar for reducing sycophancy.
GPT-5 pro
Alongside the base system, OpenAI released GPT-5 pro, which replaces o3-pro and thinks for longer using scaled parallel test-time compute. On evaluations covering over 1,000 economically valuable real-world reasoning prompts, external experts preferred GPT-5 pro over GPT-5 thinking 67.8% of the time. GPT-5 pro made 22% fewer major errors and set the family's state-of-the-art result on GPQA at 88.4% without tools.
Availability
GPT-5 is rolling out now to all Plus, Pro, Team, and Free users, with Enterprise and Edu access coming next week. The difference between free and paid tiers is usage volume. Pro subscribers get unlimited access plus GPT-5 pro; Plus users get significantly higher usage than free users. Free-tier users may wait a few days for full reasoning capabilities, and once they hit limits they transition to GPT-5 mini. Pro, Plus, and Team users can also use GPT-5 in the Codex CLI by signing in with ChatGPT.
The consolidation gives OpenAI a cleaner story for consumers and enterprise buyers alike, but the company's own caveat — that more work remains on factuality and honesty — signals the reliability race is far from settled.
Original: arxiv.org
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
118 articles
Related articles
- OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
- OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context
- OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
- OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work
- OpenAI Launches o3 and o4-mini, Its Smartest Models Yet