OpenAI Posts First Benchmark Results for Jalapeño Custom Inference Chip
OpenAI's first custom inference chip beat commercial systems on throughput per kilowatt and token latency in public benchmarks, and future generations are already underway.

Updated
Why it matters
- Jalapeño, OpenAI's first custom inference chip, delivered more peak throughput per kilowatt and lower token latency than commercial comparison systems on the InferenceX benchmark using GPT-OSS 120B.
- GPT-5.6 Sol with max reasoning reached a new high on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than another leading model.
- OpenAI's compute portfolio now includes Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank, with future Jalapeño generations already underway.
OpenAI has published the first measured performance results for Jalapeño, its first custom inference chip, and the company says it delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison. The benchmark is InferenceX, run publicly with the GPT-OSS 120B model. Jalapeño also performed strongly on DeepSeek R1 and Kimi K2, which OpenAI presents as evidence that its gains extend across model families rather than applying to a single network architecture.
The announcement arrived in a blog post titled "The full stack behind abundant intelligence," in which OpenAI laid out its compute strategy as a single integrated system spanning data centers and chips, frontier models, its developer platform, consumer and enterprise products, and AI-native devices. The company's argument for vertical integration is straightforward: each layer strengthens the next, and progress compounds fastest when the entire system improves together.
Why first-party silicon matters
Custom chips give OpenAI something it cannot buy from vendors: control over how its models run and over the economics of serving them. By developing the model, serving software, chip, memory, and network together, the company says it can improve throughput, latency, energy efficiency, and cost as one system rather than optimizing each component in isolation.
OpenAI frames Jalapeño as a "credible first-party path alongside the accelerators we use from other partners," expanding its ability to match each workload to the strongest system at the right economics. The company confirmed that future generations of the chip are already underway.
The stakes here are structural. Inference — running trained models to answer queries, power agents, and serve API customers — has become the dominant cost center in large-scale AI deployment. Every major hyperscaler has pursued custom silicon for exactly this reason: Amazon with Trainium and Inferentia, Google with its TPU line, and Meta with its MTIA accelerators. OpenAI, which rents the bulk of its compute rather than owning data centers outright, has now demonstrated working first-party silicon with measured results. That changes its negotiating position with every supplier in its portfolio.
A portfolio strategy, not a single bet
OpenAI is explicit that Jalapeño does not replace its partners. Microsoft's compute and NVIDIA's chips, the company states, "have been foundational to OpenAI's growth." The current portfolio also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank. Each brings different strengths across cloud infrastructure, accelerated computing, low-latency inference, data-center development, and energy delivery.
The company describes active management of this portfolio along two axes: capability and economics. "We use premium systems where capability matters most and optimize for efficiency where scale and cost matter more," the post states. Preserving credible choice across providers, hardware, and deployment models lets OpenAI direct demand toward the strongest performance per dollar, maintain pricing discipline as market conditions change, and move with the frontier as stronger technology emerges.
The stated goal is to stay on the Pareto frontier — the set of systems where no improvement in one dimension can be made without sacrificing another — across capability, speed, reliability, efficiency, and cost. Different workloads place different demands on the system: frontier training, high-volume inference, and always-on agents have distinct requirements across chips, software, networks, power, and latency. Different chips and providers lead on different dimensions, OpenAI notes, and that frontier keeps moving.
The operating principle the company articulates: "We partner where the ecosystem helps us move faster and build where co-design creates a meaningful advantage."
Data centers as leverage: Project Camellia
The post also highlights data-center design as another point of leverage. Project Camellia in Georgia, OpenAI says, shows how the company can design facilities around customer workloads while creating jobs, supporting local businesses, covering project infrastructure and energy costs, conserving water through a closed-loop system, and subjecting its commitments to an annual independent public audit.
The mention of independent audits is notable. AI data-center buildouts have drawn scrutiny over grid strain, water consumption, and local economic impact across the United States. OpenAI is positioning Camellia as a template that pairs infrastructure growth with verifiable community commitments.
The efficiency argument: GPT-5.6 Sol and Jevons paradox
The second half of the post makes an economic case that efficiency gains expand rather than shrink compute consumption. OpenAI cites the Artificial Analysis Coding Agent Index, where GPT-5.6 Sol with max reasoning reached a new high while using 54% fewer output tokens than another leading model.
For customers, the company says, improvements like these translate to faster results, more dependable products, fewer retries, agents that complete longer workflows, and a lower total cost for successful work. The metric OpenAI says it optimizes for is "useful intelligence per dollar" — more useful intelligence from every unit of compute.
Better models reach the right answer with fewer attempts, the post argues. Smarter routing and context management reduce wasted work. Optimized software and purpose-built hardware improve speed and energy efficiency. The value of the entire stack, in OpenAI's framing, is measured by what it produces.
OpenAI then invokes Jevons paradox, the nineteenth-century economics observation that greater efficiency in using a resource increases, rather than decreases, its total consumption. As useful intelligence becomes more capable and affordable, the company argues, more work becomes economically practical: a company can provide tailored analysis to every customer, review every contract, run live financial scenarios, and help engineers test more ideas. Greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity through more work completed, better decisions, more products launched, and more revenue generated.
The compounding argument
The post closes with OpenAI's account of its own flywheel: more productive compute and a more competitive supply base let the company serve more customers at lower cost and pass efficiency gains through to users. Growth, in turn, funds continued investment in research, infrastructure, and safety. The company calls this its "compounding advantage" — better technology creates better economics, and better economics fund the next wave of progress.
Two threads in the announcement will matter most to watch. The first is whether Jalapeño's benchmark advantage holds at production scale across OpenAI's full traffic mix, not just on public benchmarks with open-weight models. The second is how OpenAI's supplier portfolio — particularly NVIDIA, whose accelerators remain foundational to the company's growth — absorbs the emergence of a credible first-party alternative that OpenAI will presumably expand with each chip generation.
Source: OpenAI News
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
144 articles