Products & Tools

OpenAI's WebSocket Overhaul Makes Agents 40% Faster

OpenAI rebuilt the Responses API around persistent WebSocket connections, cutting agent latency by up to 40% and letting GPT-5.3-Codex-Spark hit 1,000 tokens per second in production, with bursts to 4,000 TPS.

Speeding up agentic workflows with WebSockets in the Responses API
Speeding up agentic workflows with WebSockets in the Responses APIAI-generated
By Sophie Lindqvist7 min read

Updated

Why it matters

  • Persistent WebSocket connections to the Responses API make agent workflows up to 40% faster end-to-end; GPT-5.3-Codex-Spark hit 1,000 TPS in production with bursts to 4,000 TPS.
  • A November 2025 performance sprint cut time to first token by ~45% via caching, fewer network hops, and faster safety classifiers, but structural reprocessing of conversation state remained the bottleneck.
  • Vercel's AI SDK saw latency drop up to 40%, Cline's multi-file workflows are 39% faster, and OpenAI models in Cursor became up to 30% faster after adopting WebSocket mode.

OpenAI has cut end-to-end agent latency by 40% by rebuilding the Responses API around persistent WebSocket connections, allowing its fastest coding model, GPT-5.3-Codex-Spark, to reach 1,000 tokens per second in production — with bursts up to 4,000 TPS.

The engineering shift, detailed in a technical post on OpenAI's developers blog, addresses a problem that has quietly become central to the AI industry: as GPU inference gets faster, everything around it — request validation, network hops, safety classifiers, tokenization — becomes the bottleneck. Previous flagship models like GPT-5 and GPT-5.2 ran at roughly 65 tokens per second through the Responses API. GPT-5.3-Codex-Spark, a fast coding model served on specialized Cerebras hardware optimized for LLM inference, was built to run more than an order of magnitude faster. The API's fixed overhead threatened to swallow those gains entirely.

The stakes are concrete for anyone building agentic products. When a user asks Codex to fix a bug, the agent scans the codebase for relevant files, reads them to build context, makes edits, and runs tests. Under the hood, that translates into dozens of back-and-forth Responses API requests: determine the model's next action, run a tool on the client machine, send the tool output back to the API, and repeat. Each request historically paid the full cost of reprocessing conversation state. "All of these requests can add up to minutes that users spend waiting for Codex to complete complex tasks," OpenAI writes.

When the API became the bottleneck

OpenAI breaks the latency of the Codex agent loop into three main stages: work in the API services (validating and processing requests), model inference on GPUs, and client-side time (running tools and building model context). Inference used to be the slow part, so API overhead was easy to hide. "As inference gets faster, the cumulative API overhead from an agentic rollout is much more notable," the company explains.

Around November 2025, OpenAI launched a performance sprint on the Responses API aimed at the critical-path latency of a single request. The team landed three classes of optimization: caching rendered tokens and model configuration in memory to skip expensive tokenization and network calls in multi-turn responses; eliminating network hops by removing calls to intermediate services, such as image processing resolution, and calling the inference service directly; and improving the safety stack so certain classifiers could flag conversations faster.

Those changes delivered close to a 45% improvement in time to first token (TTFT), the metric that most reflects how responsive an API feels. It still wasn't enough for GPT-5.3-Codex-Spark. Users had to wait for the CPUs running the API before they could use the GPUs serving the model.

The deeper issue was structural. OpenAI treated each Codex request as independent, processing conversation state and other reusable context on every follow-up. Even when most of a conversation hadn't changed, the system paid for work tied to the full history. As conversations grew longer, the repeated processing grew more expensive.

A persistent connection

The fix required rethinking the transport layer. Instead of establishing a new HTTP connection and sending the full conversation history with every follow-up request, the team asked whether the API could hold a persistent connection and cache state — sending only new information that required validation and processing, with reusable state held in memory for the lifetime of the connection.

OpenAI evaluated WebSockets and gRPC bidirectional streaming, and landed on WebSockets. As a simple message transport protocol, it meant users wouldn't have to change their Responses API input and output shapes. It was developer-friendly and fit OpenAI's existing architecture with little disruption.

The first prototype changed what the team believed was possible for Responses API latency. An engineer on the Codex team with deep expertise across the API stack built it by running a Codex agent overnight.

That prototype modeled an entire agentic rollout as a single long-running response. Using asyncio, the Responses API would asynchronously block in the sampling loop after the model sampled a tool call, then send a response.done event to the client. After the client executed the tool, it sent back a response.append event with the tool result, unblocking the sampling loop and letting the model continue.

The analogy OpenAI uses is treating a local tool call like a hosted one. When the model calls web search, the inference loop blocks, calls the search service, and puts the response into model context. The prototype did the same thing — except instead of calling a remote service, it sent the model's tool call to the client over the WebSocket, then inserted the client's tool result into the context and continued sampling.

The design eliminated repeated API work across an agent rollout: preinference work happened once, the loop paused for tool execution, and postinference work happened once at the end. The cost was an unfamiliar, more complicated API shape. OpenAI wanted developers to adopt WebSocket support without rewriting their integrations around a new interaction mode.

Familiar API, incremental stack

The shipped version reverted to a familiar shape. Developers keep using response.create with the same body, and use previous_response_id to continue conversation context from a previous response's state.

On a WebSocket connection, the server maintains a connection-scoped, in-memory cache of previous response state. When a follow-up response.create includes previous_response_id, the API fetches that state from the cache instead of rebuilding the full conversation. The cached state includes the previous response object, prior input and output items, tool definitions and namespaces, and reusable sampling artifacts such as previously rendered tokens.

That cache enabled several downstream optimizations. Safety classifiers and request validators now process only new input rather than the full history on every turn. Rendered tokens are cached and appended to, skipping unnecessary tokenization. Successful model resolution and routing logic is reused across requests. Non-blocking postinference work, such as billing, overlaps with subsequent requests.

The goal, OpenAI writes, was to get as close as possible to the minimal-overhead prototype while keeping an API shape developers already understood and had built around.

Production results

After a two-month sprint building WebSocket mode, OpenAI launched an alpha with key coding agent startups so they could integrate it into their infrastructure and safely ramp up traffic before general availability.

The launch results were immediate. Codex moved the majority of its Responses API traffic onto WebSocket mode and saw significant latency improvements. For GPT-5.3-Codex-Spark, OpenAI hit its 1,000 TPS target and observed bursts up to 4,000 TPS, demonstrating that the API could keep up with dramatically faster inference under real production load.

The gains propagated quickly through the developer ecosystem:

  • Vercel integrated WebSocket mode into the AI SDK and reported latency decreases of up to 40%.
  • Cline's multi-file workflows are 39% faster.
  • OpenAI models in Cursor became up to 30% faster.
  • Codex users running the latest models, including GPT-5.3-Codex and GPT-5.4, all benefit from WebSocket mode's speed-up.

Alpha users, OpenAI notes, "loved it," reporting up to 40% improvements in their agentic workflows before the public launch.

Why it matters beyond OpenAI

WebSocket mode ranks among the most significant new capabilities in the Responses API since its launch in March 2025, and it arrived fast: OpenAI went from idea to production in a few weeks through close collaboration between its API and Codex teams.

The broader signal for the industry is architectural. Coding agents and other agentic products live and die by perceived responsiveness, and the economics of specialized inference hardware from vendors like Cerebras are collapsing generation times. That only transfers to users if the surrounding software — transport protocols, caching layers, safety classifiers, routing — scales at the same pace. OpenAI frames the lesson directly: "as model inference gets faster, the services and systems that surround inference also need to speed up to transfer these gains to users." For competitors building agent platforms, the Responses API's WebSocket redesign now sets a reference point for what production agent infrastructure needs to look like.

Original: x.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

114 articles

Related articles

  1. OpenAI Ships GPT-5.3-Codex-Spark, a Coding Model Built for Real-Time Work
  2. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  3. OpenAI unveils GPT-5-Codex, a coding-tuned variant of GPT-5
  4. OpenAI ships GPT-5.3-Codex, its first self-built coding model
  5. OpenAI's GPT-5.1-Codex-Max Flags Coming Cybersecurity Threshold

« Previous articleNext article »