Models

Google's Gemini 4 Argon Catches GPT-6 Astra, but Claude Opus 5.5 Stays Ahead

Gemini 4 Argon matches GPT-6 Astra in independent testing but trails Claude Opus 5.5 — and burns over twice the tokens per task as Astra despite low per-token pricing.

By Rebecca Stone4 min read

Updated

Why it matters

  • Gemini 4 Argon is Google's first frontier model in over seven months.
  • Independent testing shows Argon matches GPT-6 Astra but trails Anthropic's Claude Opus 5.5.
  • Argon's per-token price is low, but it uses more than twice as many tokens per task as GPT-6 Astra.

Google's Gemini 4 Argon matches OpenAI's GPT-6 Astra in independent testing but falls short of Anthropic's Claude Opus 5.5, according to the benchmark results reported by The Decoder. The release ends a frontier-model drought at Google that lasted more than seven months.

The stakes here are straightforward. Frontier AI models are the products on which Google, OpenAI, and Anthropic are building their enterprise API businesses, their consumer subscriptions, and their cloud margins. When a lab goes more than half a year without a new frontier release while its two closest rivals ship, perception shifts: customers evaluating multi-year platform commitments start asking whether the leader board has settled without them. Gemini 4 Argon is Google's answer to that question, and the answer is a qualified one.

What the benchmarks show

Independent testing puts Gemini 4 Argon on par with GPT-6 Astra. That is a meaningful result for Google, because it closes the gap with OpenAI at the top of the market. It does not, however, put Google in front. Anthropic's Claude Opus 5.5 outperforms Argon in the same testing, which means Anthropic retains the strongest model position and Google settles for parity with OpenAI rather than a clear lead of its own.

For buyers, parity at the top tier changes the calculus from "which model is best" to "which model is cheapest and most reliable for my workload." That framing matters for the next part of the story, because Argon's headline pricing advantage does not survive contact with its usage patterns.

The token economics problem

The Decoder reports that Gemini 4 Argon's per-token price is low. On paper, that positions it as the budget option among frontier models. In practice, the model consumes more than twice as many tokens per task as GPT-6 Astra. Do the arithmetic: a per-token rate cut in half is erased entirely if the model needs two tokens for every one a competitor uses. At an overhead above twofold, the effective cost per completed task is higher than Astra's despite the lower sticker price.

This is the kind of detail that enterprise procurement teams will catch quickly. API bills are computed on tokens consumed, not on advertised rates, and agentic workloads — where a model reasons through long chains of steps — amplify token consumption on every cycle. A model that thinks verbosely is a model that costs more at scale, whatever its per-token price says. Token efficiency has become a benchmark in its own right, and on that measure Argon trails by a wide margin.

For Google, the efficiency gap also carries an infrastructure cost. Every extra token generated consumes inference compute. At frontier-model scale, a twofold token overhead per task translates into materially more accelerator time per customer workload. That pressures the unit economics of Google's API business even before price competition does.

A staggered rollout

Gemini 4 Argon is reaching select testers first. The API and paid tiers follow later, according to The Decoder. The pattern is a familiar one across the industry: labs put new frontier models in the hands of trusted external evaluators and early partners before exposing them to general API traffic, partly to manage capacity and partly to surface failure modes before they reach paying customers.

The staggered release also means the competitive picture today is provisional. Independent testers have the model; production users do not yet. Real-world usage at scale tends to differ from benchmark results, and token-consumption figures observed in testing may shift as Google tunes the model. Buyers making platform decisions in the next few months will want to re-run their own cost comparisons once the API opens.

Why this matters beyond Google

The result tightens a three-way race. Anthropic holds the performance lead with Claude Opus 5.5. OpenAI and Google now sit effectively level on capability, with GPT-6 Astra holding a significant edge in token efficiency over Argon. That configuration rewards Anthropic on quality, OpenAI on cost-per-task, and leaves Google competing on the argument that parity plus low per-token pricing equals value — an argument the efficiency numbers undercut.

There is also a research signal in the data worth naming plainly: raw capability and token efficiency have decoupled. Argon matches Astra on benchmarks while using more than twice the tokens, which tells you the two labs are trading off inference compute against performance in different ways. Whether Google can close the efficiency gap in future versions, rather than buying benchmark parity with brute token volume, will determine whether Argon's parity is durable or expensive theater.

The next concrete milestone is the public API launch. When it arrives, third-party developers can verify the token-consumption figures on their own workloads — and that independent accounting, more than any leaderboard, will decide where Gemini 4 Argon actually lands in the market.

(This article is based on reporting by The Decoder.)

Original: the-decoder.de

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

170 articles

Related articles

  1. Google Announces Gemini 4 Argon, But No One Outside Can Use It
  2. Google Launches Gemini 4 Argon, Its Most Advanced AI Model Yet
  3. Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
  4. Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
  5. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month

« Previous article