Models

OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price

OpenAI's GPT-5.1 hits the API with adaptive reasoning, a no-reasoning mode, 24-hour prompt caching, and new shell and apply_patch tools — all at GPT-5 prices.

Introducing GPT-5.1 for developers
Introducing GPT-5.1 for developershackNY / Openverse
By Sophie Lindqvist6 min read

Updated

Why it matters

  • GPT-5.1 reaches 76.3% on SWE-bench Verified (all 500 problems), up from GPT-5's 72.8%, at identical pricing and rate limits.
  • Balyasny Asset Management reports GPT-5.1 runs 2-3x faster than GPT-5 while using about half the tokens of leading competitors on tool-heavy reasoning tasks.
  • New features include a 'no reasoning' mode (reasoning_effort='none', now the default), prompt caching extended to 24 hours with cached tokens 90% cheaper, and new apply_patch and shell tools in the Responses API.

OpenAI has released GPT-5.1 in its API platform, a new model in the GPT-5 series that the company says dynamically adjusts how much time it spends thinking based on task complexity, delivering significantly faster and more token-efficient responses on everyday work. The release targets the two workloads where developer money is currently flowing fastest: agentic systems and coding. Pricing and rate limits remain identical to GPT-5.

The core engineering change is what OpenAI calls adaptive reasoning. On straightforward tasks, GPT-5.1 spends fewer tokens thinking, which translates into snappier products and lower inference bills. On hard problems, the model stays persistent, exploring options and checking its work to preserve reliability. OpenAI offers a concrete example: asked "show an npm command to list globally installed packages," GPT-5.1 answers in 2 seconds instead of 10.

Early testers put numbers on the claim. Balyasny Asset Management said GPT-5.1 "outperformed both GPT-4.1 and GPT-5 in our full dynamic evaluation suite, while running 2-3x faster than GPT-5," and that across tool-heavy reasoning tasks it "consistently used about half as many tokens as leading competitors at similar or better quality." Pace, an AI insurance BPO, reported its agents run "50% faster on GPT-5.1 while exceeding accuracy of GPT-5 and other leading models across our evals."

A new "no reasoning" mode

Alongside adaptive reasoning, GPT-5.1 introduces a genuine "no reasoning" mode. Developers set reasoning_effort to 'none', and the model behaves like a non-reasoning model for latency-sensitive use cases while retaining frontier intelligence and performant tool-calling. OpenAI says that relative to GPT-5 with 'minimal' reasoning, the no-reasoning configuration is better at parallel tool calling, coding, instruction following, and using search tools — and it supports web search in the API platform.

Sierra measured the difference in production: GPT-5.1 on "no reasoning" mode showed a "20% improvement on low-latency tool calling performance compared to GPT-5 minimal reasoning" in its real-world evals.

The mode ships as the default. GPT-5.1 now defaults to 'none', which OpenAI calls ideal for latency-sensitive workloads. The company recommends 'low' or 'medium' for higher-complexity tasks and 'high' when intelligence and reliability matter more than speed.

24-hour prompt caching

OpenAI also extended prompt caching from a few minutes to as long as 24 hours of cache retention. Longer-lived caches mean more follow-up requests can reuse cached context, cutting latency and cost for long-running interactions such as multi-turn chat, coding sessions, and knowledge retrieval workflows. The economics are notable: cached input tokens remain 90% cheaper than uncached tokens, with no additional charge for cache writes or storage. Developers activate it with the parameter "prompt_cache_retention='24h'" on the Responses or Chat Completions API. Priority Processing customers will also see noticeably faster performance with GPT-5.1 over GPT-5.

Coding gains, benchmarked

OpenAI developed GPT-5.1's coding behavior with direct feedback from Cursor, Cognition, Augment Code, Factory, and Warp. The result, according to the company: a more steerable coding personality, less overthinking, improved code quality, better user-facing update messages during sequences of tool calls, and more functional frontend designs — especially at low reasoning effort.

On SWE-bench Verified, where a model receives a code repository and issue description and must generate a working patch, GPT-5.1 reaches 76.3% across all 500 problems, up from GPT-5's 72.8%, and OpenAI notes it works even longer than its predecessor on difficult problems.

The customer quotes stack up. Augment Code called GPT-5.1 "more deliberate with fewer wasted actions, more efficient reasoning, and better task focus," reporting "more accurate changes, smoother pull requests, and faster iteration across multi-file projects." Cline said GPT-5.1 "achieved SOTA on our diff editing benchmark with a 7% improvement, demonstrating exceptional reliability for complex coding tasks." CodeRabbit called it its "top model of choice for PR reviews." Cognition said it is "noticeably better at understanding what you're asking for and working with you to get it done." Factory cited "noticeably snappier responses" and reasoning depth that adapts to the task. Warp is making GPT-5.1 the default for new users, saying it "builds on the impressive intelligence gains that the GPT-5 series introduced, while being a far more responsive model."

JetBrains offered the most emphatic endorsement. "GPT 5.1 isn't just another LLM—it's genuinely agentic, the most naturally autonomous model I've ever tested. It writes like you, codes like you, effortlessly follows complex instructions, and excels in front-end tasks, fitting neatly into your existing codebase," said Denis Shiryaev, Head of AI DevTools Ecosystem at JetBrains.

Two new tools: apply_patch and shell

GPT-5.1 arrives with two new tools in the Responses API. The freeform apply_patch tool lets the model create, update, and delete files using structured diffs — no JSON escaping required. Instead of suggesting edits, the model emits patch operations that an application applies and reports back on, enabling iterative, multi-step editing workflows. Developers enable it with "tools": [{"type": "apply_patch"}].

The shell tool goes further: the model proposes shell commands, the developer's integration executes them on a local machine, and outputs return to the model. OpenAI describes this as a simple plan-execute loop that lets models inspect the system, run utilities, and gather data until a task is complete. A model that can write and run its own shell commands is a meaningful expansion of API-level autonomy — and a security decision developers will have to own in their integrations.

Availability and the numbers that matter

GPT-5.1 and gpt-5.1-chat-latest are available on all paid API tiers at GPT-5 pricing. OpenAI also released gpt-5.1-codex and gpt-5.1-codex-mini, optimized for long-running agentic coding tasks in Codex or Codex-like harnesses. The company does not currently plan to deprecate GPT-5 in the API and promises advance notice if that changes.

Beyond SWE-bench Verified, the reported evaluation deltas are mostly positive but not uniform. GPT-5.1 (high) scores 88.1% on GPQA Diamond without tools versus 85.7% for GPT-5, 85.4% on MMMU versus 84.2%, 26.7% on FrontierMath with a Python tool versus 26.3%, and 67.0% on Tau 2-bench Airline versus 62.6%. It trails on AIME 2025 (94.0% versus 94.6%), Tau 2-bench Telecom (95.6% versus 96.7%, with OpenAI noting it gave GPT-5.1 a short generically helpful prompt there) and Tau 2-bench Retail (77.9% versus 81.1%). BrowseComp Long Context at 128k is flat at 90.0%.

The stakes are straightforward. Reasoning models changed the cost equation by burning tokens on thinking, and every provider is now racing to make that thinking conditional. OpenAI's bet with GPT-5.1 is that efficiency — fewer tokens on easy tasks, longer persistence on hard ones, cheaper caching — matters as much to developers as raw benchmark scores. The company says developers can expect more capable agentic and coding models "in the weeks and months ahead."

Original: platform.openai.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

114 articles

Related articles

  1. OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
  2. OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
  3. OpenAI's GPT-5.5 System Card Details Safety Push and Pro Variant
  4. OpenAI unveils GPT-5-Codex, a coding-tuned variant of GPT-5
  5. OpenAI's WebSocket Overhaul Makes Agents 40% Faster

« Previous articleNext article »