Google Announces Gemini 4 Argon, But No One Outside Can Use It
Google claims its new Gemini 4 Argon model beats GPT-6 Astra and Opus 5.5 on coding benchmarks — but only Google engineers can use it so far.

Updated
Why it matters
- Google announced Gemini 4 Argon with industry-leading claims in coding, knowledge work, and cybersecurity, but external users cannot access it yet.
- Gemini 4 Argon scores 77.9 percent on the DeepSWE v1.1 benchmark, higher than GPT-6 Astra, Fable 5.1, and Opus 5.5, per Google.
- Google says Argon agents migrated more than 800,000 lines of the Fuchsia OS Zircon kernel from C/C++ to Rust and helped save 300 TiB of memory across its data centers.
Google has announced Gemini 4 Argon, a new AI model it claims delivers industry-leading performance in coding, knowledge work, and cybersecurity — and no one outside the company can use it yet.
The announcement comes after a summer in which Google promised Gemini 3.5 Pro in June but instead released smaller, faster Flash models while saying 3.5 Pro was still in testing. With Argon, Google signals a return to the frontier model race against rivals like OpenAI's GPT-6 Astra, Fable 5.1, and Anthropic's Opus 5.5.
Internal deployment, external silence
While external users wait, Google says its own engineers are already using the model extensively. According to the company, Argon used "fleet-wide telemetry data" to help Google save 300 TiB of memory across its data centers.
Argon agents are also migrating C/C++ codebases to Rust throughout Google. The company reports the effort has touched thousands of lines in the core re2 and libgav1 libraries and more than 800,000 lines in the Fuchsia OS Zircon kernel.
The internal-only launch is notable because it reverses the usual cadence of frontier releases, where labs typically ship models to developers and customers alongside benchmark announcements. Google is instead using its own infrastructure as the proving ground.
Benchmark claims
Google backs its claims with a set of benchmark numbers. On the software engineering DeepSWE v1.1 benchmark, Gemini 4 Argon scores 77.9 percent, which the company says is higher than GPT-6 Astra, Fable 5.1, and Opus 5.5.
Google also promises similar strength across a range of long-horizon tasks, pointing to Argon's industry-leading score on the Vals Index, an economic analysis test.
These numbers, if they hold up under independent testing, would put Google back at the top of the coding model leaderboard after months focused on cheaper, faster Flash variants. The claims remain unverified until the model reaches outside hands.
Why it matters
The frontier race has concrete stakes for developers and enterprises choosing model providers, and Google's delayed Gemini 3.5 Pro created an opening for competitors. Argon's internal Rust migration work — more than 800,000 lines converted in the Zircon kernel alone — also points to where Google sees agentic AI heading: long-running, large-scale code transformation rather than single-turn chat.
Google has not said when Gemini 4 Argon will be available to external users, and it has not addressed the status of the still-unreleased Gemini 3.5 Pro. Whenever the model does ship, expect independent benchmarkers to test whether the 77.9 percent DeepSWE score survives contact with real-world workloads.
Original: blog.google
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
167 articles
Related articles
- Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
- Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
- Google DeepMind chief says Gemini 4 is nearly ready for launch
- Google Launches Gemini 3 Pro at $2/Million Input Tokens
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs