OpenAI Ships GPT-5.3-Codex-Spark, a Coding Model Built for Real-Time Work
GPT-5.3-Codex-Spark, the first model from the OpenAI–Cerebras partnership, generates over 1,000 tokens per second and is rolling out to ChatGPT Pro users as a research preview.

Updated
Why it matters
- GPT-5.3-Codex-Spark is OpenAI's first model designed for real-time coding, delivering over 1,000 tokens per second on Cerebras' Wafer Scale Engine 3.
- Pipeline optimizations cut roundtrip overhead by 80%, per-token overhead by 30%, and time-to-first-token by 50% across the serving stack.
- The model launches as a research preview for ChatGPT Pro users, text-only with a 128k context window, and is the first in a planned family of ultra-fast models.
OpenAI has released GPT-5.3-Codex-Spark, a smaller version of GPT-5.3-Codex and the company's first model designed specifically for real-time coding, delivering more than 1,000 tokens per second on Cerebras' ultra-low latency hardware. The model is available today as a research preview for ChatGPT Pro users in the latest versions of the Codex app, CLI, and VS Code extension.
The release marks the first concrete milestone of the partnership OpenAI announced with Cerebras in January. Codex-Spark runs on Cerebras' Wafer Scale Engine 3, a purpose-built AI accelerator for high-speed inference that gives Codex what OpenAI calls a "latency-first serving tier." The stakes here are structural: inference speed is emerging as a competitive axis in the coding-assistant market, where the difference between waiting seconds and seeing results instantly changes how developers interact with models at all.
Two Modes of Work
OpenAI's framing positions Codex-Spark as a complement to, not a replacement for, its frontier models. The company's latest frontier models excel at long-running tasks, working autonomously for hours, days or weeks without intervention. Codex-Spark targets the opposite end of the spectrum: making targeted edits, reshaping logic, or refining interfaces and seeing results immediately. With Codex-Spark, the company says, Codex now supports both long-running, ambitious tasks and getting work done in the moment.
The interactive design runs deep. Developers can collaborate with the model in real time, interrupting or redirecting it as it works, and rapidly iterate with near-instant responses. Because the model is tuned for speed, it keeps a lightweight default working style: it makes minimal, targeted edits and does not automatically run tests unless asked.
Benchmark Performance and Speed
OpenAI reports that Codex-Spark demonstrates strong performance on SWE-Bench Pro and Terminal-Bench 2.0, two benchmarks evaluating agentic software engineering capability, while accomplishing tasks in a fraction of the time compared to GPT-5.3-Codex. The company did not publish specific benchmark scores in the announcement.
According to OpenAI, task duration is estimated as the sum of output generation time (output tokens divided by sampling speed), prefill time, total tool execution time, and total network overhead. The speed advantage comes from both raw sampling throughput on Cerebras silicon and from pipeline optimizations OpenAI made elsewhere.
Latency Improvements Across the Fleet
Training Codex-Spark surfaced a broader lesson: model speed alone was not enough for real-time collaboration. "We also needed to reduce latency across the full request-response pipeline," OpenAI writes. The company implemented end-to-end latency improvements in its harness that will benefit all models.
The engineering changes are specific and measurable. OpenAI streamlined how responses stream from client to server and back, rewrote key pieces of its inference stack, and reworked session initialization so the first visible token appears sooner. Through a persistent WebSocket connection and targeted optimizations inside the Responses API, the company reduced overhead per client/server roundtrip by 80%, per-token overhead by 30%, and time-to-first-token by 50%. The WebSocket path is enabled for Codex-Spark by default and will become the default for all models soon.
Those numbers matter beyond this single release. If these pipeline gains carry across OpenAI's fleet, every model served through the stack gets faster without any change to the underlying weights.
The Cerebras Strategy
OpenAI is explicit that Cerebras is a specialized tool, not a replacement for GPUs. GPUs, the company says, remain foundational across its training and inference pipelines and deliver the most cost-effective tokens for broad usage. Cerebras complements that foundation by excelling at workflows that demand extremely low latency, tightening the end-to-end loop so Codex feels more responsive as developers iterate. The two can also be combined for single workloads to reach the best performance.
Sean Lie, CTO and Co-Founder of Cerebras, framed the preview as an opening move rather than a finished product:
"What excites us most about GPT-5.3-Codex-Spark is partnering with OpenAI and the developer community to discover what fast inference makes possible—new interaction patterns, new use cases, and a fundamentally different model experience. This preview is just the beginning."
The partnership also carries market weight for Cerebras itself. Serving an OpenAI production workload on Wafer Scale Engine 3 hardware is a validation the chipmaker can point to as it competes against Nvidia's dominant GPU position in inference.
Availability and Constraints
Access comes with conditions. Because Codex-Spark runs on specialized low-latency hardware, usage is governed by a separate rate limit that may adjust based on demand during the research preview. Usage will not count toward standard rate limits, but OpenAI warns that when demand is high, users may see limited access or temporary queuing as the company balances reliability across users. OpenAI says it is working with Cerebras to ramp up datacenter capacity, harden the end-to-end user experience, and deploy larger frontier models on the same infrastructure.
At launch, Codex-Spark has a 128k context window and is text-only. OpenAI is also making the model available in the API for a small set of design partners to learn how developers want to integrate it into their products, with expanded access planned over the coming weeks as the integration is tuned under real workloads.
Safety Evaluation
OpenAI states that Codex-Spark includes the same safety training as its mainline models, including cyber-relevant training. The company evaluated the model as part of its standard deployment process, which includes baseline evaluations for cyber and other capabilities, and determined that it does not have a plausible chance of reaching its Preparedness Framework threshold for high capability in cybersecurity or biology.
What Comes Next
Codex-Spark is the first in a planned family of ultra-fast models. OpenAI says it will introduce larger models, longer context lengths, and multimodal input as it learns where fast models shine for coding.
The longer-term vision is a Codex with two complementary modes that blend over time: longer-horizon reasoning and execution on one side, real-time collaboration for rapid iteration on the other. In practice, OpenAI envisions Codex keeping a developer in a tight interactive loop while delegating longer-running work to sub-agents in the background, or fanning out tasks to many models in parallel when breadth and speed matter — so developers don't have to choose a single mode up front.
The underlying bet is straightforward. As models become more capable, interaction speed becomes a clear bottleneck. Ultra-fast inference tightens that loop, and OpenAI is positioning the combination of Cerebras hardware and its rewritten serving stack as the way to remove it — with GPT-5.3-Codex-Spark as the first proof point.
Original: cerebras.ai
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles