Products & Tools

OpenAI cuts GPT-6 agent costs with smarter prompt caching

OpenAI's GPT-6 prompt caching system offers discounts up to 90% on cached tokens, with a new dashboard, miss diagnostics, and prewarming to cut persistent-agent costs.

By James Calloway5 min read

Updated

Why it matters

  • OpenAI offers discounts of up to 90% on cached input tokens for GPT-6.
  • Cache discounts apply to eligible shared prefixes reused within a 30-minute window.
  • A new Prompt Caching Dashboard tracks cache hit rates and cached versus uncached token composition.
  • The diagnostics tool pinpoints cache misses; OpenAI's example shows 5,629 tokens lost to a tools_changed miss.
  • Developers can now change reasoning effort between responses on GPT-6 without breaking cache.

OpenAI now offers discounts of up to 90% on cached input tokens for GPT-6, and the company says an improved prompt caching system delivers higher cache hit rates by default across the model family. The change targets the fastest-growing category of AI workloads: persistent agents that run for hours on code refactoring, research documents, and presentations.

The upgrade arrives as agent-based applications reshape how developers consume API capacity. These applications issue long sequences of API requests that build on one another, repeatedly carrying forward the same instructions, tool definitions, and context from earlier turns. Caching that shared context lets OpenAI reuse computation across requests, cutting response times alongside the input-token discount.

What changed with GPT-6 prompt caching?

According to OpenAI's announcement, the new system gives cache discounts for eligible shared prefixes reused within a 30-minute window. That window defines how long cached prefixes remain eligible for reuse before a fresh request must reprocess the input from scratch.

Three capabilities ship alongside the default performance improvements:

  • A Prompt Caching Dashboard that shows how much of an application's input is served from cache, tracks hit rates over time, and breaks down cached versus uncached tokens in an input composition chart.
  • A prompt caching diagnostics tool that explains unexpected cache misses by comparing a request against a recent response to identify changes to the model, tools, settings, or input that blocked reuse.
  • Explicit cache breakpoints that let developers choose which prompt prefixes to cache, documented in a refreshed prompt caching guide.

Why does cache visibility matter for agent developers?

Until now, developers had limited ability to understand why a cache hit failed. A single silent change — a reordered tool definition, an altered setting — could wipe out reuse across an entire session and inflate costs without any obvious signal.

The diagnostics tool addresses that gap directly. When a miss occurs, it reports the reason, the comparison's reusable tokens, and the missed token count. OpenAI's own example shows the shape of the output:

{
  "prompt_cache_diagnostics": {
    "type": "cache_miss",
    "reason": "tools_changed",
    "comparison_reusable_tokens": 5629,
    "cache_missed_tokens": 5629
  }
}

In that example, 5,629 tokens lost their discount because the tools changed between requests. The estimated number of affected tokens, OpenAI says, helps developers "assess the size of the impact and decide how you can optimize your integration to maximize cache hit rates." That per-miss quantification turns caching from a black box into a tunable engineering discipline.

The dashboard serves the complementary monitoring role. Its views help developers spot drops in cache hits over time and evaluate how application changes affect caching performance — a feedback loop that matters most for agents whose context grows across dozens of turns.

Can you change reasoning effort without losing the cache?

Yes, and OpenAI framed this as one of the most practical additions. On GPT-6 models, developers can now change reasoning effort between responses without breaking cache. Raising effort for a hard task or lowering it for a routine follow-up is done by appending a configuration_update while leaving request-level reasoning effort unchanged.

The distinction matters for agent economics. Previously, adjusting how much reasoning a task needed could invalidate the reusable context an agent had accumulated. Now developers can tune compute per turn while preserving the cached prefix, which OpenAI says "lets you adjust how much reasoning a task needs while preserving reusable context."

How do you keep cache intact as tools and instructions evolve?

OpenAI's guidance for agents whose tool needs shift over a session centers on stability:

  • Keep tool definitions, schemas, and ordering stable so earlier context stays reusable.
  • Use allowed_tools to expose only relevant callable tools, or set tool_choice to none when no tools are needed, instead of removing definitions outright.
  • Use new developer messages to append instructions toward the end of the context, overriding older ones rather than rewriting them from the top.

The pattern is consistent: mutate the tail of the context, never the head. Because caching depends on shared prefixes, any change near the front of the prompt invalidates everything after it — the reason tool ordering carries the same weight as tool content.

What about latency?

OpenAI also introduced prewarming, which prepares known context ahead of time so the model can start responding sooner when a request arrives. An application can prewarm shared instructions, tool definitions, or reference material during startup, before the user asks a first question. As the company puts it, this "moves processing out of the user's wait time" — a latency win that compounds over a long agent session rather than a single query.

What is the practical takeaway?

OpenAI describes all the new controls — breakpoints, diagnostics guidance, tool-stability practices, prewarming — as optional layers on top of the engine's default caching performance, letting developers tailor caching to their workload. For teams running agents on GPT-6, the operational checklist from the announcement is straightforward: monitor hit rates in the Prompt Caching Dashboard, investigate unexpected misses with the diagnostics tool, and follow the prompt caching guide to improve the setup — or use Codex to review code, apply improvements, and measure results.

The release signals where OpenAI sees API economics heading. As agentic workloads stretch from single-turn requests to multi-hour sessions, input caching is becoming a first-class cost lever rather than an incidental optimization, and the up-to-90% discount on cached tokens makes cache hit rates a metric that directly shapes the unit economics of persistent agents.

Source: OpenAI News

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

200 articles

Related articles

  1. OpenAI Launches GPT-6.1 Sol at One-Fifth of Astra's Price
  2. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  3. OpenAI's WebSocket Overhaul Makes Agents 40% Faster
  4. OpenAI Ships GPT-6.1 Sol at a Fifth of Astra's Price
  5. OpenAI ships GPT-5.4 mini and nano for coding, tool use, and agent workloads

« Previous articleNext article »