Chips & Compute

Agentic AI Is Driving a CPU Comeback — and a Shortage

Agentic AI is pushing tool calls, tokenization and safety checks back onto CPUs. AWS is rationing cycles, Intel is sold out, and AMD doubled its server CPU forecast.

The CPU Comeback Is Upon Us
The CPU Comeback Is Upon UsPeterTea / Openverse
By James Calloway5 min read

Updated

Why it matters

  • AWS told engineers to conserve CPU cycles at all costs amid exploding wait times for CPU server capacity, The Information reported.
  • AMD's Madhu Rangarajan: "In our testing, seven of the eight stages in realistic agentic AI pipelines run entirely on the CPU."
  • Intel has sold out of server CPUs through at least the end of the year; AMD has doubled its server CPU forecast; Nvidia unveiled Vera, an Arm-based CPU for agents.

Amazon Web Services told its engineers earlier this year to conserve CPU cycles at all costs, after wait times for CPU server capacity exploded as AI workloads strained the company's cloud infrastructure, according to a report in The Information.

The directive caught many off guard. The AI boom drove surging demand for GPUs and, later, memory. CPUs sat out that story because their relative lack of parallelization made them a poor fit for running and serving large language models. Agentic AI is changing that. Agentic systems allow AI models to operate autonomously and call on sub-agents, and each step in those pipelines lands heavily on the CPU.

Matt Kimball, vice president and principal data center analyst at Moor Insights & Strategy in Austin, Texas, says 2026 has brought a spike in CPU demand, much of it driven by agentic AI. "It's one thing to have this agentic workload, and let's say it spawns 100 agents. If I'm going to roll this out across my enterprise, those 100 become tens of thousands, hundreds of thousands, or millions of agents," Kimball says. "You have agents spawning sub-agents, making API calls and talking to more agents through [Anthropic's] model context protocol."

Agents need computers, and computers need CPUs

The problem centers on "tool use" — an LLM's ability to access the internet, open files, and operate software to accomplish its task. The LLM's inference still runs primarily on a GPU, but the tool calls it makes typically execute on the CPU.

"Many components of an agentic AI task are inherently CPU based jobs," explains Souvik Kundu, senior staff research scientist at Intel. "The CPU does the job of parsing output, figuring out which tool to invoke, making the API call or running the code, collecting the result, and feeding it back." Madhu Rangarajan, vice president of compute and enterprise AI products at AMD, puts a number on it: "In our testing, seven of the eight stages in realistic agentic AI pipelines run entirely on the CPU."

Kundu co-authored a paper on agentic AI optimization with researchers at Georgia Tech in Atlanta. The group found the CPU often sits idle while inference runs on the GPU, and the GPU idles while tool calls run on the CPU. Their proposed scheduling optimizations cut end-to-end latency by up to 1.8 times under sustained load.

The gains chase a moving target. Agentic systems generate work at machine speed and multiply it as they go. During OpenAI's inadvertent hack of Hugging Face, its model fired off as many as 300 actions an hour, and a single agent can spawn sub-agents that make tool calls of their own.

Safety guardrails add further CPU load. Kundu says policy checks often inspect syntax and log files, and guardrails may use small models — under a billion parameters — to analyze task complexity or intent. Because those models are small and latency-sensitive, the work tends to stay on the CPU rather than move to a GPU.

Tokenization adds to bottlenecks

A second paper, co-authored by Georgia Tech PhD student Euijun Chung, found that servers with too few CPU cores fall behind on dispatching work to GPUs, causing the GPUs to stall while waiting for instructions.

The paper also identifies tokenization as a pressure point. Tokenization converts text into integer token IDs the model can process, and unlike the matrix math that dominates LLM inference, it is branchy, data-dependent sequential string manipulation. Tokenizing small prompts is trivial. Agentic models are not small-prompt workloads: they must parse and tokenize the result of every tool call.

"If you have an ongoing sequence of, say, 100,000 tokens, and you have a tool result of 1,000 tokens, the tokenizer will have to tokenize the whole sequence again. And you have to do tokenization at every agentic tool call," Chung says.

The paper finds that time-to-first-token latency can increase dramatically as sequence length grows, and that more CPU cores reduce the problem. In test runs at longer sequence lengths, increasing CPU core counts cut time-to-first-token latency by roughly 1.5 to 7 times. Chung and his colleagues tested smaller models — Alibaba's Qwen 3-30B and Meta's Llama 3.1-70B — due to hardware limits. He speculates larger models will face less dramatic bottlenecks because of their higher overall GPU demand, but expects agentic AI to push token lengths far beyond what the team tested. "If you think about something like Anthropic's Claude, you can easily hit 500,000, even a million tokens," Chung says. "In the world of agentic AI, the average sequence length will grow and grow, so I'm expecting this problem to get worse in future workloads."

The crunch may just be starting

The market signals back up the research. Intel has sold out of server CPUs through at least the end of the year. AMD has doubled its server CPU forecast. Arm and Qualcomm have announced new CPUs designed to accelerate agentic AI. Nvidia has prioritized Vera, its Arm-based CPU for agents, as part of its Vera Rubin platform.

Kimball calls the surge in demand an "absolute tell" that CPUs are now considered a key part of an agentic AI system. He warns it may translate into broader CPU shortages and higher prices, much as already occurred with GPUs and memory. "You're already seeing a CPU crunch to some degree. When you look at the constraints in the market, it even trickles down into the consumer space," Kimball says, pointing to Intel's decision to cut client CPU production in favor of server CPUs despite growing client sales on its new 18A process. CPU makers, in his view, will follow the money.

Original: aws.amazon.com

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

119 articles

Related articles

  1. OpenAI Models Are Coming to Amazon Bedrock via a Stateful Agent Runtime
  2. Cloudflare Brings OpenAI's GPT-5.4 to Agent Cloud for Enterprises

« Previous articleNext article »