Google's Gemini 4 'Carbon' Model Reportedly Matches Anthropic's Opus 5.5 on Coding
A Google employee told Business Insider that an unreleased Gemini 4 variant called Carbon matches Anthropic's Opus 5.5 on coding, as new modes appear in the Gemini app ahead of the wider launch.
Updated
Why it matters
- An anonymous Google employee told Business Insider that an unreleased Gemini 4 variant codenamed Carbon matches Anthropic's Opus 5.5 on coding tasks
- Gemini 4 'Argon' is the lower tier currently in limited early rollout, with no general public availability
- New modes have appeared inside the Gemini app and AI Studio the same week as the Carbon report, signaling a wider launch
- Google has not publicly confirmed the existence of the Carbon tier
- Business Insider did not publish a specific benchmark score, transcript, or independent test backing the comparison
A Google employee has privately told Business Insider that an unreleased Gemini 4 variant codenamed "Carbon" performs at parity with Anthropic's Opus 5.5 on coding tasks. The comparison surfaced even as the prior "Argon" tier remains unavailable to most users.
The Decoder surfaced the Business Insider report this week. The single hard claim is narrow: one Google staffer, cited anonymously, said Carbon's coding performance matches Opus 5.5. Business Insider did not publish a benchmark score, a side-by-side transcript, or an independent test. The comparison is informal, not a formal Google announcement.
The name Carbon sits above Argon in Google's apparent codename ladder for the Gemini 4 generation. The arrangement suggests Google is reserving the highest tier for an internal or limited rollout, similar to the "Ultra" tier Google has used in past cycles.
What does the leak actually say?
The claim is narrow but pointed. A single employee, quoted anonymously, told Business Insider that Carbon's coding output rivals Anthropic's most advanced publicly referenced model, Opus 5.5. Business Insider relayed the remark. The Decoder then surfaced it.
That kind of internal brag typically leaks when a launch is close. Anthropic itself has used similar comparisons in its own model release materials. Google has historically flagged coding as the metric that decides which lab leads the frontier.
The leak matters because coding has become the most-watched capability in the frontier-model race. Anthropic's Opus line has held a strong position with developers. Coding evaluations also drive enterprise purchasing decisions more directly than chat-quality tests, since coding results are scored against test suites that pass or fail.
A Google model that genuinely ties Opus 5.5 on programming tasks would be the first credible external claim of parity from a major competitor. It would also undercut the framing that Anthropic has been running ahead on the metric developers care most about.
Where does Gemini 4 actually stand today?
Google has begun seeding Gemini 4 "Argon" through limited channels but has not pushed it to the general public. The slow rollout follows the pattern Google has used for previous Gemini generations.
Staged access has historically let Google manage capacity. It has also let the company gather feedback from enterprise customers first, and run a quiet safety review before public deployment.
Carbon, the tier above Argon, has not been confirmed by Google outside the employee remark. Google's public communications have not mentioned Carbon by name. The lack of an on-the-record statement is itself a signal: Google rarely pre-announces top-tier models before they ship.
What changed in the Gemini app?
The same week the Carbon report surfaced, new modes appeared inside the Gemini app and the AI Studio developer console. The Decoder reported that these modes suggest Google is preparing the wider Gemini 4 release.
The new options point at reasoning depth, tool use, and code execution. Exact labels vary by account. Mode additions are the most concrete public signal that the launch window is opening.
Google has used similar interface changes ahead of past Gemini drops. Internal codenames on public surfaces have historically appeared within a short window of flagship releases. The coincidence of the new app modes and the Carbon report points to a launch inside a single news cycle.
What does "matching" actually mean?
The Business Insider source did not specify which coding benchmark Carbon ties on. Anthropic has published results on SWE-bench, HumanEval, Aider, and several internal evaluations across Opus generations.
Without a specific benchmark, "matches" is a soft claim. It could mean ties on SWE-bench Verified, parity on a long-horizon agent evaluation, or simply equivalent output quality on internal tests Google has not disclosed.
That ambiguity cuts both ways. It lets Google claim parity without committing to a specific score. It also lets Anthropic dismiss the claim until a number lands publicly.
The benchmark conversation will only settle once Google releases Carbon and third-party testers run their own evaluations. Until then, the only public reading of the claim is that Google believes it is competitive on coding, not that it has demonstrated parity on any specific test.
Why does the Anthropic comparison matter?
Anthropic's Opus line has set the ceiling other labs chase. Earlier Opus releases scored at or near the top of public coding benchmarks. Opus 5.5 extended that lead into the current generation, and Anthropic has marketed the result aggressively to enterprise buyers.
Coding has become the proxy metric enterprise buyers use to compare frontier models. It produces measurable output. It now drives agent products that handle real workflows, from refactoring codebases to running pull-request reviews.
Google has invested heavily in coding capability across earlier Gemini versions, but Anthropic has held the mindshare position with developers. A Google model that genuinely matches Opus 5.5 on coding would give enterprise buyers a credible alternative. It would also reset the conversation about which lab leads on programming.
How does this fit Google's launch pattern?
Google has used tiered codenames before. The Gemini 1.x cycle included Pro, Ultra, and Nano variants. Gemini 2 and 3 cycles extended the naming into brands like Flash and Pro. Gemini 4 appears to use Argon and Carbon as internal labels, while public-facing branding may still use Pro or Ultra.
The internal-versus-public naming split is a recurring pattern in frontier-model launches. It lets labs reference performance internally without committing to specific product names until pricing, capacity, and rollout timing are locked.
The Argon-and-Carbon framing suggests Google is preparing at least two tiers within Gemini 4. The lower tier, Argon, is already in early rollout. The higher tier, Carbon, remains internal.
What happens next?
Google has not commented on Carbon. The company typically declines to discuss unreleased products by codename. Employees have signed non-disclosure agreements before.
The Argon rollout will likely continue through enterprise channels first. Broader public availability will likely depend on capacity rather than a fixed calendar. Google has historically run public launches in waves, with early access for AI Studio users first, then Gemini app users, then API customers.
Business Insider's report will put pressure on Google to either confirm or deny the Carbon framing. Anthropic, which has not been named in Google's public communications on Gemini 4, will be watching the rollout closely because coding parity is the line it cannot afford to cede.
The next data point will be public benchmarks. Until Google ships Carbon or Argon widely, the only signal is internal: one employee, one comparison, one leaked remark. The launch window implied by the new app modes will determine whether the claim holds up under outside scrutiny.
If the claim holds, the competitive balance inside enterprise coding tools shifts. If it does not, the leak becomes a footnote in another round of the frontier-model race.
Original: x.com
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
225 articles
Related articles
- Google's Gemini 4 Argon Catches GPT-6 Astra, but Claude Opus 5.5 Stays Ahead
- Google Announces Gemini 4 Argon, But No One Outside Can Use It
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
- Google Unveils Gemini 4 Argon — but Only 'Trusted Cyber Defenders' Get It First
- Google DeepMind chief says Gemini 4 is nearly ready for launch