Models

OpenAI ships GPT-5 to developers at $1.25 per million input tokens

OpenAI released GPT-5 to its API in three sizes at $1.25 per million input tokens, scoring 74.9% on SWE-bench Verified and 96.7% on τ2-bench telecom with new verbosity and custom tools parameters.

By Rebecca Stone6 min read

Updated

Why it matters

  • GPT-5 launched in the OpenAI API at $1.25 per million input tokens and $10 per million output tokens, with gpt-5-mini at $0.25/$2 and gpt-5-nano at $0.05/$0.40.
  • GPT-5 scored 74.9% on SWE-bench Verified (up from o3's 69.1%) and 88% on Aider polyglot, while using 22% fewer output tokens and 45% fewer tool calls than o3 at high reasoning effort.
  • GPT-5 hit 96.7% on τ2-bench telecom, a benchmark no model scored above 49% on when Sierra.ai published it two months ago.
  • GPT-5 accepts 272,000 input tokens and emits up to 128,000 tokens for a 400,000-token total context window, and scores 95.2% on OpenAI-MRCR at 128K.
  • GPT-5 makes roughly 80% fewer factual errors than o3 on LongFact and FActScore prompts, with a FActScore hallucination rate of 2.8% versus o3's 23.5%.

OpenAI released GPT-5 to its API platform on Thursday, pricing the flagship model at $1.25 per million input tokens and $10 per million output tokens and posting state-of-the-art results on the coding and agentic benchmarks it prioritizes for developer use.

On SWE-bench Verified, a real-world software engineering evaluation, GPT-5 scored 74.9%, up from o3's 69.1%. On Aider polyglot, a code-editing benchmark, GPT-5 set a new record of 88%, cutting error rates by roughly one-third compared to o3. On τ2-bench telecom, a tool-calling benchmark Sierra.ai published two months ago, GPT-5 hit 96.7%; no model in the original paper crossed 49%.

What does the release change for developers?

GPT-5 arrives in three sizes — gpt-5, gpt-5-mini, and gpt-5-nano — and on the Responses API, the Chat Completions API, and as the default in Codex CLI. Pricing scales down sharply with the smaller variants:

  • gpt-5: $1.25 / 1M input tokens, $10 / 1M output tokens
  • gpt-5-mini: $0.25 / 1M input tokens, $2 / 1M output tokens
  • gpt-5-nano: $0.05 / 1M input tokens, $0.40 / 1M output tokens

The API version is the reasoning model that powers maximum performance in ChatGPT, not the routing system that decides between fast and reasoning mode in the consumer product. A separate non-reasoning endpoint, gpt-5-chat-latest, mirrors the $1.25/$10 pricing.

GPT-5 also runs across Microsoft 365 Copilot, Copilot, GitHub Copilot, and Azure AI Foundry on day one.

How does GPT-5 improve on o3?

OpenAI published a benchmark suite that frames GPT-5 as a broad upgrade over its predecessor, not just a coding specialist. Highlights from the company's tables:

  • AIME 2025 (no tools): 94.6% vs. o3's 88.9%
  • GPQA diamond (no tools): 85.7% vs. o3's 83.3%
  • FrontierMath with Python: 26.3% vs. o3's 15.8%
  • HMMT 2025: 93.3% vs. o3's 81.7%
  • MMMU multimodal: 84.2% vs. o3's 82.9%
  • Scale MultiChallenge (graded by o3-mini): 69.6% vs. o3's 60.4%

On SWE-Lancer's IC SWE Diamond freelance coding tasks, GPT-5 completed $112,000 worth of the benchmark's work, compared with $86,000 for o3.

GPT-5 also reaches its SWE-bench Verified score with greater efficiency. Relative to o3 at high reasoning effort, the new model uses 22% fewer output tokens and 45% fewer tool calls.

What new knobs does the API get?

Three additions target the developer experience directly:

  • A verbosity parameter (low, medium, high) that controls default answer length, with medium as the default. Explicit length instructions still override the parameter.
  • A minimal value on reasoning_effort that returns answers faster by skipping extended thinking, alongside the existing low, medium, and high values.
  • Custom tools, a new tool type that lets GPT-5 call tools with plaintext instead of JSON. Developers can constrain the format with a regex or a context-free grammar.

OpenAI designed custom tools around an escaping problem. JSON tool calls require models to correctly escape quotation marks, backslashes, newlines, and other control characters. On long inputs like hundreds of lines of code or a 5-page report, that error rate creeps up. Plaintext calls sidestep the issue. GPT-5 scores about the same on SWE-bench Verified using custom tools.

GPT-5 also supports parallel tool calling, the built-in web search, file search, and image generation tools, streaming, Structured Outputs, prompt caching, and the Batch API.

What do the alpha testers say?

OpenAI shared quotes from six partner companies that ran GPT-5 on private evaluations before launch.

Cursor called GPT-5 "the smartest model we've used" and "remarkably intelligent, easy to steer, and even has a personality we haven't seen in other models." Windsurf said GPT-5 hit state-of-the-art on its evals and "has half the tool calling error rate over other frontier models." Vercel said "it's the best frontend AI model, hitting top performance across both the aesthetic sense and the code quality, putting it in a category of its own."

Manus reported GPT-5 "achieved the best performance we've ever seen from a single model on our internal benchmarks." Notion said "the model's rapid responses, especially in low reasoning mode, make GPT-5 an ideal model when you need complex tasks solved in one shot." Inditex added that "what truly sets GPT-5 apart is the depth of its reasoning: nuanced, multi-layered answers that reflect real subject-matter understanding."

In OpenAI's own head-to-head comparisons with o3 on frontend web development, human testers preferred GPT-5 70% of the time.

How does long-context performance look?

GPT-5 accepts a maximum of 272,000 input tokens and emits up to 128,000 reasoning and output tokens, for a 400,000-token total context. OpenAI highlighted OpenAI-MRCR, its long-context information retrieval benchmark:

  • 2-needle at 128K: 95.2% vs. o3's 55.0%
  • 2-needle at 256K: 86.8% (o3 was not reported at this length)

On BrowseComp Long Context, a new long-context Q&A benchmark OpenAI open-sourced alongside the release, GPT-5 answered correctly 90% of the time at 128K inputs and 88.8% at 256K inputs.

What about hallucinations and safety?

OpenAI leaned into the agentic-correctness angle. On LongFact and FActScore prompts, the company said GPT-5 makes roughly 80% fewer factual errors than o3. The benchmark table backs that up: GPT-5's FActScore hallucination rate is 2.8% versus o3's 23.5%.

GPT-5 also extends across instruction-following. On COLLIE it scored 99.0%. On Scale MultiChallenge, graded by o3-mini, it hit 69.6%. On OpenAI's internal hard instruction-following eval, it scored 64.0% versus o3's 47.4%.

OpenAI flagged that the default MultiChallenge grader (GPT-4o) frequently mis-scores responses, and that swapping to o3-mini as grader improved accuracy on samples they inspected.

Where does this leave the competitive landscape?

GPT-5's launch comes as the major AI providers have shifted their messaging from raw capability to agentic reliability. OpenAI's framing — long-horizon tool use, fewer errors, steerable verbosity, plaintext tool calls — maps onto the pain points enterprise developers cite when moving from demos to production.

The pricing also reflects that shift. The flagship model's $10 per million output tokens matches o3's pricing. The aggressive drop to $0.40 per million output tokens on gpt-5-nano gives OpenAI a low-latency option aimed at high-volume workloads where reasoning overhead hurts more than it helps.

For agentic coding products specifically, the GPT-5 API release sets up a direct comparison with Anthropic's Claude and Google's Gemini on the same SWE-bench Verified, Aider polyglot, and τ2-bench telecom benchmarks. Cursor, Windsurf, GitHub Copilot, and Codex CLI all list GPT-5 as a first-class option on day one.

Whether GPT-5's lead on those benchmarks translates into shipped product wins will depend on what developers build on top of it over the next quarter.

Source: OpenAI News

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

214 articles

Related articles

  1. OpenAI ships GPT-5.4 mini and nano for coding, tool use, and agent workloads
  2. OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context
  3. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  4. OpenAI Launches GPT-6.1 Sol at One-Fifth of Astra's Price
  5. First Look at GPT-5: Leading Developers Test OpenAI's New Model

« Previous articleNext article »