Chips & Compute

OpenAI's Jalapeño Chip: LLMs Cut Design Time to Record Lows

OpenAI's first accelerator went from concept to silicon in under 20 months with fewer than 100 engineers, using internal LLMs to write RTL and benchmark code — and cutting latency up to 3.6x versus Nvidia's GB300.

How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
How OpenAI Used Its Own LLMs to Design Its Jalapeño Chipseanrnicholson / Openverse
By Sophie Lindqvist4 min read

Updated

Why it matters

  • Jalapeño delivers up to 13.4 petaflops of 4-bit compute and cut end-to-end latency by up to 3.6x versus Nvidia's GB300 in OpenAI-cited benchmarks, while using less power.
  • The chip moved from concept to first silicon in under 20 months, with nine months from first RTL to tape-out, using an in-house team averaging fewer than 100 people plus Broadcom for physical design.
  • AI models lifted performance on DeepSeek's multi-head latent attention kernel benchmark from 0.31 percent to 88.94 percent of the theoretical ceiling in roughly 40 hours after first silicon returned in May.

OpenAI's debut AI accelerator, Jalapeño, delivers up to 13.4 petaflops of 4-bit compute and cut end-to-end inference latency by up to 3.6 times compared to Nvidia's GB300 in benchmarks the company cited at its 25 August unveiling. The chip, now entering OpenAI's inference fleet, pairs 232 gigabytes of the most advanced memory available with 15.4 terabytes per second of bandwidth, and consumes less power than the Nvidia hardware OpenAI currently relies on.

The benchmark numbers may or may not survive contact with production workloads. The design timeline is the part of the story with broader stakes: Jalapeño went from first architecture concept to first silicon in under 20 months, with only nine months separating the first register-transfer level (RTL) code from tape-out. OpenAI achieved that with an in-house team averaging fewer than 100 people, according to Richard Ho, vice president of hardware at OpenAI. That figure spans system design, software, and supply chain roles, but not the engineers at Broadcom, which handled physical design "from the gates onward."

"The models are giving superpowers to our engineers," Ho says. "Our engineers are still driving the work. They're still the final arbiter of what's going on. But they can do things a lot faster. They can explore a lot more paths."

Experts outside the company see the timeline as significant, with caveats. Andrew Kahng, distinguished professor at the University of California, San Diego, calls the speed "likely best in class today." David Chin, co-founder of agentic chip design startup Verkor.io, says "the schedule they gave us is quite credible" but believes Broadcom's involvement was essential: "If you have somebody else start from scratch, it won't be possible." His co-founder Ravi Krishna calls it "a relatively impressive result" and expects that today's improved LLMs could compress the schedule further.

Why LLMs fit the front end

The technical logic behind OpenAI's workflow is straightforward. "Automation itself has existed in chip design for many decades. It's not a new problem," says Ankur Srivastava, director of semiconductor initiative and innovation at the University of Maryland. What distinguishes LLMs is their grasp of language and code, which suits tasks "still in the linguistic domain of the problem."

OpenAI's front-end workflow was built around Accelerated Hardware Synthesis (XLS), an open-source high-level synthesis toolchain originally developed at Google. Engineers wrote in DSLX, a domain-specific language inspired by Rust, and C++; XLS converted the code to Verilog. "We were thinking about how to leverage AI to make the project faster, and the AI was much better at software-looking things," says Chris Leary, member of technical staff at OpenAI. "XLS in some ways looks like software, so it got that benefit." Leary started the XLS project during his time at Google. Kahng agrees the approach "has legs" going forward.

The AI assistance extended past design into bring-up. When the first chips returned from the foundry in May, the team pointed internal models at writing benchmark software. On DeepSeek's multi-head latent attention kernel benchmark, performance climbed from 0.31 percent of the theoretical ceiling to 88.94 percent in roughly 40 hours. Ho says the result is repeatable and reshapes planning: "All our schedule assumptions are going to be based on the fact we have this capability now."

The models themselves evolved mid-project. Work began with OpenAI's o3, released publicly in April 2025. By the end, the team had precursors to GPT-6 Astra, which was not publicly released until 3 September 2026. Leary says the newer model works directly in Verilog without XLS's translation layer and is close to operating proprietary design tools on its own. Ho also confirmed the team used internal LLMs fine-tuned for chip design and declined to detail them, but said lessons will reach commercial models: "It's safe to say that Astra and following models will be very good at chip design."

Backend gains, with limits

At IEEE Hot Chips 2026, Ho and Leary quantified AI-guided physical design optimization, including a 10 percent area reduction for the matrix multiplication units against an optimized human baseline. Verkor's founders say the backend approach already looks conservative, an artifact of the project's October 2024 start. "From April [2026] onwards…is when they really started to be able to handle those tasks better," Ravi Krishna says. Co-founder Suresh Krishna adds that "there's no reason you couldn't have an agentic loop that largely accelerates the backend of the process as well."

Ho says second and third-generation designs, already underway with the team at roughly 100 people, will introduce AI in verification and physical design, aided by new tools for automatic waveform analysis during debugging. Both engineers draw a firm line at full autonomy. "We're not saying that anyone can come and just build state-of-the-art, frontier AI/ML accelerator chips using just Codex," Ho says. "We are saying some very specific things about how to be better at Codex and how we are focusing on a small team and fast timelines to reach quality results."

Original: inferencex.semianalysis.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

115 articles

Related articles

  1. OpenAI's Jalapeño Chip Beats Rivals on Power and Latency
  2. OpenAI and Broadcom Unveil Jalapeño, a Custom LLM Inference Chip
  3. OpenAI Commits to 6 Gigawatts of AMD GPUs in Multi-Year Deal
  4. AWS and OpenAI Sign $38 Billion Multi-Year Compute Deal

« Previous articleNext article »