Models

OpenAI ships GPT-4.1 family in API with 1M-token context window

OpenAI ships GPT-4.1 family in the API with a 1M-token context window and a 54.6% SWE-bench Verified score, pricing nano at $0.10 per million input tokens and retiring GPT-4.5 on July 14, 2025.

Introducing GPT-4.1 in the API
Introducing GPT-4.1 in the APIAI-generated
By Sophie Lindqvist6 min read

Updated

Why it matters

  • GPT-4.1 scores 54.6% on SWE-bench Verified, a 21.4-point absolute gain over GPT-4o
  • Context window expands to 1 million tokens, up from 128,000 for GPT-4o
  • GPT-4.1 nano is priced at $0.10 per million input tokens; GPT-4.1 mini cuts cost by 83% versus GPT-4o
  • GPT-4.5 Preview will be retired from the API on July 14, 2025
  • Knowledge cutoff for the GPT-4.1 family is June 2024

OpenAI began rolling out the GPT-4.1 family in its API this week, posting a 54.6% score on the SWE-bench Verified coding benchmark and supporting 1 million-token context windows across three new models priced as low as $0.10 per million input tokens.

The flagship GPT-4.1 cleared 54.6% of SWE-bench Verified tasks, a 21.4-percentage-point jump over GPT-4o and a 26.6-point lift over GPT-4.5, OpenAI said. The company framed the release as a pragmatic step rather than a frontier-science breakthrough. "While benchmarks provide valuable insights, we trained these models with a focus on real-world utility," OpenAI wrote in its announcement.

The new family consists of GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano — the company's first nano-tier offering. All three carry a refreshed knowledge cutoff of June 2024 and process up to 1 million tokens of context, up from 128,000 for prior GPT-4o models.

What does GPT-4.1 actually beat GPT-4o at?

OpenAI's internal numbers put GPT-4.1 ahead of GPT-4o across three categories that drive most paid API usage: coding, instruction following, and long-context retrieval.

On Scale's MultiChallenge benchmark, which tests multi-turn instruction following, GPT-4.1 scored 38.3%, a 10.5-point absolute improvement over GPT-4o. On IFEval, GPT-4.1 hit 87.4% versus 81.0% for GPT-4o. The model also set a new state-of-the-art on Video-MME's long, no-subtitles category at 72.0%, a 6.7-point gain.

Coding showed the widest spread. GPT-4.1 more than doubled GPT-4o's score on Aider's polyglot diff benchmark, and beat GPT-4.5 by 8 points absolute on the same test. Extraneous edits on internal code evaluations dropped from 9% with GPT-4o to 2% with GPT-4.1, according to OpenAI.

Paid human graders preferred GPT-4.1's frontend code over GPT-4o's 80% of the time in head-to-head comparisons, the company said.

How big is the context jump?

The 1 million-token window is roughly eight copies of the full React codebase, OpenAI said — large enough to ingest most enterprise code repositories or stacks of legal documents in a single prompt. The previous GPT-4o ceiling was 128,000 tokens.

GPT-4.1 maintained needle-in-a-haystack accuracy across all positions up to the full 1 million tokens in OpenAI's internal test. The company open-sourced a harder benchmark, OpenAI-MRCR, that requires models to disambiguate between multiple identical requests hidden in context. GPT-4.1 hit 57.2% on the two-needle variant at 128k tokens, versus 31.9% for GPT-4o.

On a second new dataset called Graphwalks, which requires multi-hop breadth-first search across a context filled with hexadecimal hash graphs, GPT-4.1 scored 61.7% at 128k tokens — matching OpenAI's o1 reasoning model and beating GPT-4o handily.

What did alpha testers see?

Six enterprise partners tested GPT-4.1 ahead of launch. The results skewed strongly in OpenAI's favor:

  • Windsurf: 60% higher score on its internal coding benchmark; 30% more efficient tool calling; 50% less likely to make unnecessary edits
  • Qodo: produced the better suggestion in 55% of cases across 200 real GitHub pull requests
  • Blue J: 53% more accurate on hard tax scenarios
  • Hex: nearly 2× improvement on the most challenging SQL evaluation
  • Thomson Reuters: 17% better multi-document review accuracy for its CoCounsel legal assistant
  • Carlyle: 50% better retrieval on very large dense documents; first model to overcome lost-in-the-middle errors

What's the pricing story?

GPT-4.1 costs 26% less than GPT-4o on median queries, OpenAI said. The full per-million-token table:

Model Input Cached input Output
gpt-4.1 $2.00 $0.50 $8.00
gpt-4.1-mini $0.40 $0.10 $1.60
gpt-4.1-nano $0.10 $0.025 $0.40

Prompt caching discount rises to 75% for the new models, up from 50% previously. Long-context requests carry no premium beyond standard per-token pricing. The Batch API adds another 50% discount.

GPT-4.1 mini matches or exceeds GPT-4o on most intelligence evaluations while cutting latency by nearly half and cost by 83%. GPT-4.1 nano is OpenAI's fastest and cheapest model to date. "It delivers exceptional performance at a small size," the company wrote, noting 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding.

What happens to GPT-4.5?

OpenAI will retire GPT-4.5 Preview from the API on July 14, 2025, three months from the deprecation announcement. The company introduced GPT-4.5 in February as a research preview to explore a compute-intensive architecture.

"GPT-4.1 offers improved or similar performance on many key capabilities at much lower cost and latency," OpenAI wrote. The company said it would carry forward GPT-4.5's creativity, writing quality, humor, and nuance into future API models — language that reads as a soft acknowledgment that 4.5 was a costly experiment that didn't earn its keep at scale.

GPT-4.1 will not ship in ChatGPT. "Many of the improvements in instruction following, coding, and intelligence have been gradually incorporated into the latest version of GPT-4o," OpenAI said, with more coming in future ChatGPT releases.

Why does this matter for the agent build-out?

OpenAI positioned the family as infrastructure for agentic systems. The company pointed to the Responses API as the orchestration primitive and listed four flagship use cases: real-world software engineering, extracting insights from large documents, resolving customer requests with minimal hand-holding, and other "complex tasks."

Better instruction following — the internal hard-prompt eval rose from 29.2% for GPT-4o to 49.1% for GPT-4.1 — directly addresses the most common complaint from agent developers: models that drift from system prompts or ignore tool-calling instructions. Function-calling results were mixed. GPT-4.1 scored 65.5% on ComplexFuncBench, just below GPT-4o's 66.5%, but cleared 49.4% on Taubench airline and 68.0% on Taubench retail.

What does this signal about OpenAI's roadmap?

The release carves a clearer line between OpenAI's reasoning models (o1, o3-mini) and its "workhorse" GPT series. GPT-4.1 does not reason explicitly the way o1 does, yet on SWE-bench Verified it outperforms OpenAI's own o1 (41.0%) and approaches o3-mini high (49.3%). For coding agents specifically, GPT-4.1 may become the default backbone — fast, cheap, and reliable enough to drive multi-step tool use.

The nano tier also opens a new market segment for OpenAI: latency-sensitive, classification-style workloads where price per token matters more than raw intelligence. At $0.10 per million input tokens, GPT-4.1 nano undercuts most third-party small models on a pure-cost basis.

The July 14 retirement of GPT-4.5 closes one chapter and opens another: developers who relied on 4.5 for writing quality now have three months to migrate, with no clearly designated successor carrying that banner in the 4.1 lineup.

Original: scale.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

209 articles

Related articles

  1. OpenAI ships GPT-5 to developers at $1.25 per million input tokens
  2. OpenAI ships GPT-5.4 mini and nano for coding, tool use, and agent workloads
  3. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  4. OpenAI Ships GPT-6.1 Sol at a Fifth of Astra's Price
  5. OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context

« Previous article