OpenAI's Jalapeño Chip Beats Rivals on Power and Latency
OpenAI's first custom inference chip delivered 1.5–1.9× more AI work per watt and up to 3.6× lower latency on InferenceX benchmarks, with deployment planned by year-end.

Updated
Why it matters
- Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency than comparison systems across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T on the SemiAnalysis InferenceX benchmark.
- The chip is rated at 700 watts but measured sustained power stayed at or below 550 watts on tested workloads; AI helped compress design to tapeout in nine months.
- OpenAI plans to deploy Jalapeño in its compute infrastructure by end of year, with Gen 2 deep in development and Gen 3 taking shape; it will continue deploying NVIDIA and other partners' accelerators.
OpenAI's first custom inference chip, Jalapeño, delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than leading commercially available AI systems in the company's first published benchmarks. The results, measured on SemiAnalysis's public InferenceX benchmark across three open-weight models, mark OpenAI's entry into first-party silicon with numbers the company claims place the chip on the Pareto frontier for both power efficiency and latency.
The stakes are considerable. Inference — serving responses to users — is where AI companies spend the bulk of their operating cost at scale, and power efficiency translates directly into margin and capacity. By producing more useful work from the same power and hardware, OpenAI says Jalapeño can help it serve more demand and lower the cost of delivering a successful result. For OpenAI, that can improve operating leverage by allowing useful work and revenue to grow faster than the cost to serve.
The benchmark numbers
OpenAI tested Jalapeño on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request, comparing it against leading commercially available systems across the tested operating range — from high-throughput serving to highly interactive, low-latency use. The evaluation covers GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, models developed both inside and outside OpenAI.
Although performance is often reported per chip, OpenAI argues the more useful standard is performance per unit of power. To normalize results, the company used each accelerator's published chip power rating. Jalapeño is rated at 700 watts, but its measured sustained power remained at or below 550 watts on the workloads tested.
The headline figures:
- GPT-OSS 120B: approximately 1.9× higher peak mixed tokens per second per kilowatt (85,448 vs. 44,960 mixed/kW), 1.7× lower end-to-end latency (1.03 s vs. 1.80 s), 2.7× lower minimum time-between-tokens (0.69 vs. 1.87 ms, corresponding to 1,459 vs. 535 tok/s/user), and 53.7× more throughput at the comparison system's previous time-between-tokens (22,935 vs. 427 mixed/kW at 535.28 tok/s/user).
- DeepSeek R1 670B: approximately 1.7× higher peak mixed/kW (19,641 vs. 11,781), 3.6× lower end-to-end latency (1.65 s vs. 5.99 s), 4.1× lower minimum time-between-tokens (1.43 vs. 5.90 ms, 700 vs. 169 tok/s/user), and 104.3× more throughput at previous time-between-tokens (12,258 vs. 118 mixed/kW at 169.41 tok/s/user).
- Kimi K2.5 1T: approximately 1.5× higher peak mixed/kW (18,195 vs. 11,862), 3.4× lower end-to-end latency (1.56 s vs. 5.31 s), 3.8× lower minimum time-between-tokens (1.44 vs. 5.48 ms, 694 vs. 182 tok/s/user), and 56.1× more throughput at previous time-between-tokens (6,744 vs. 120 mixed/kW at 182.46 tok/s/user).
The dramatic "throughput at previous TBT" multiples — 53.7×, 104.3×, and 56.1× — describe how much more aggregate work Jalapeño sustains while holding per-user response speed at the level the comparison systems could only achieve at far lower utilization. In internal testing, OpenAI says Jalapeño's advantage widened further on frontier OpenAI models, suggesting the architecture becomes more valuable as workloads grow larger and more demanding.
Throughput and latency without the tradeoff
Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two. For customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows.
OpenAI evaluates performance at a matched user experience, measuring how much useful AI work each system can complete per unit of power while meeting the latency that customers and interactive agents require. The company frames this as especially important for agents, which need to complete many steps in sequence, so delays can compound across an entire task.
Why the architecture matters
Language-model inference moves through several distinct phases with different bottlenecks. Prefill, when the system processes a prompt, is compute-intensive. Decode, when the system generates the response token by token, is constrained more by memory bandwidth. Communication can also add latency when data must move between cores and chips, leaving some processing units idle while they wait. A system that excels at one phase can lose that advantage while waiting for data or moving model state between different resources.
OpenAI designed Jalapeño to minimize data movement and communication delays. Model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase. The network is integral to the architecture: its large domain allows the entire workload to remain within one connected system, keeping the complete request fast and efficient from beginning to end.
The result, OpenAI says, is a balanced and fungible accelerator that can support changing model architectures, excel at both prefill and decode, and adapt as the balance between them changes — a defining feature of agentic workloads.
AI built the chip, and AI programs it
Jalapeño's development itself demonstrates a full-stack advantage. AI enabled the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip's arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule. Earlier generations of OpenAI models helped the team design and bring up the chip, while the latest models are accelerating how OpenAI optimizes and programs it.
The chip was designed as a clear, predictable programming target for both humans and AI. Engineers describe work through local tensors, explicit communication, and predictable synchronization; AI then optimizes how that work is mapped, placed, scheduled, and coordinated across the system. OpenAI says that clear, predictable structure gives AI a tractable way to tackle the traditionally difficult problem of parallel programming.
Supporting each new model family still requires new kernels and model-specific optimizations. Using Codex with GPT-Astra, the team brought three open-weight models that were not part of Jalapeño's original production plan to high performance within two months — demonstrating, in OpenAI's telling, both the flexibility of the architecture and the speed at which AI can help program it. For selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than the existing human-expert-written implementations. OpenAI cautions those figures apply to the selected blocks, not the full model, but frames them as pointing toward a powerful new development loop.
Deployment timeline and what comes next
OpenAI plans to begin deploying Jalapeño within its compute infrastructure by the end of the year. The company describes it as the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will build on what OpenAI learns and further advance both efficiency and speed.
The company is not abandoning its suppliers. Meeting growing demand for AI will require more compute from every available source, and OpenAI says it will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads. Before deployment, the company is continuing production qualification, maturing the software, preparing to operate Jalapeño at scale, and validating performance across more models.
In practical terms, OpenAI says Jalapeño expands what is possible for efficient, low-latency inference in three ways: ultra-fast-mode inference at efficiencies previously available only in fast mode, fast-mode inference at efficiencies previously available only in batched mode, and higher efficiency for batched-mode inference.
The strategic significance extends beyond one chip. Jalapeño is working first-party silicon with measured results, and OpenAI frames it as evidence that designing models, products, serving software, chips, memory, networking, and systems together — using what the company learns from real workloads to improve every layer of the stack — yields advantages no single-layer vendor can easily match. If the AI-assisted development loop that compressed Jalapeño's design-to-tapeout cycle to nine months holds for Gen 2 and Gen 3, the gap between OpenAI's full-stack approach and commodity inference infrastructure could widen with each generation.
Source: OpenAI News
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles