OpenAI and Broadcom Unveil Jalapeño, a Custom LLM Inference Chip
OpenAI and Broadcom have introduced Jalapeño, a custom AI chip built for LLM inference, targeting better performance, efficiency and scale across AI systems as inference costs shape AI economics.

Updated
Why it matters
- OpenAI and Broadcom introduced Jalapeño, a custom AI chip built for LLM inference.
- The companies say the chip improves performance, efficiency, and scale across AI systems.
- The announcement did not include specifications, benchmarks, pricing or deployment timelines.
OpenAI and Broadcom have introduced Jalapeño, a custom AI chip built specifically for large language model inference. The companies say the chip is designed to improve performance, efficiency, and scale across AI systems.
The announcement puts a name on a partnership that pairs OpenAI's inference workload demands with Broadcom's track record in custom silicon design. Jalapeño targets the inference side of the AI stack — the stage where a trained model generates outputs for users — rather than training. That distinction matters commercially: inference now accounts for the bulk of day-to-day AI compute as chatbots, coding assistants and API traffic grow.
The chip's stated goals, according to the announcement, are threefold: better performance, better efficiency, and the ability to scale across AI systems. Inference-optimized silicon is a growing category because general-purpose GPUs are not tailored to the memory-bandwidth and latency profile of token generation. Chips tuned for that workload can lower the cost per query, which is a central constraint on AI services economics.
For OpenAI, custom silicon is a way to reduce dependence on merchant GPU suppliers as demand for its models grows. For Broadcom, the design win extends its business building application-specific chips for large AI customers. Broadcom has positioned itself as a partner for companies that want dedicated accelerators without building a full chip design organization in-house.
The companies describe Jalapeño as built for LLM inference to improve "performance, efficiency, and scale across AI systems." The announcement does not detail specifications, benchmarks, pricing, or deployment timelines.
The launch lands amid a broader industry shift toward inference-focused hardware. Model providers are under pressure to cut the cost of serving models at scale, and dedicated inference chips are one of the main levers available. OpenAI's move mirrors a wider pattern among frontier AI companies that are commissioning or acquiring custom accelerator designs to complement commodity hardware.
The stakes extend beyond the two companies. If inference-optimized chips meaningfully lower serving costs, they change the unit economics of AI products — from consumer subscriptions to enterprise API pricing. They also reshape demand across the semiconductor supply chain, shifting some volume from general-purpose accelerators toward bespoke designs.
What remains open is execution: Jalapeño's impact will depend on when it ships, how it performs against incumbent hardware in real inference workloads, and how widely OpenAI deploys it across its systems. Those details will determine whether the chip becomes a meaningful cost lever or a niche engineering exercise.
Source: OpenAI News
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
121 articles