Google Announces Gemini 4 Argon, Its Next Frontier Model
Google's Gemini 4 Argon ships first to trusted cyber defenders without cyber guardrails, with 1M-token output, 77.9% on DeepSWE, and $2/$10 per million token pricing under a phased release.
Updated
Why it matters
- Gemini 4 Argon rolls out first to trusted cyber defenders via the Fairwind Program, released to them without cyber guardrails; pricing is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off.
- Argon expands output capacity to 1M tokens (up from 64K) and posts 77.9% on DeepSWE v1.1, 51.3% and #1 on Zapier's AutomationBench, 91.7% on LVBench, and a tied-first 68% on CWE-bench v1.
- Internal Google deployments include a 40% beat on a published quantum optimization baseline, 300+ TiB of memory freed across data centers, and Argon agents migrating up to 800K+ lines of C/C++ to Rust, with a libgav1 Rust decoder running 2.7x faster.
Google has announced Gemini 4 Argon, a new frontier model built to sustain deep reasoning across complex, long-horizon workflows, and it is rolling the model out first to a set of trusted cyber defenders through its Fairwind Program. The company says Argon "is fundamentally changing the way we work and build at Google" and delivers frontier performance in real-world software engineering, enterprise knowledge work including legal and finance, and cybersecurity defense.
The stakes are considerable. Argon arrives with offensive-capable cybersecurity skills — Google says it will release the model to trusted defenders and its own internal teams "without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities." That capability profile, combined with the phased release under the U.S. government's voluntary process for pre-release model access, makes Argon a test case for how frontier models with dual-use security capabilities reach the market.
Pricing and availability
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input token price. That undercuts the effective cost of long-horizon work dramatically, and long-horizon work is exactly what Argon is built for: Google is expanding the model's output token limit to an industry-leading 1M tokens, up from the previous 64K tokens.
"When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go," Google said in its announcement.
Broad availability is not immediate. Google says it is "actively engaged in the U.S. government's voluntary process for pre-release model access" and will gather feedback from early testers as it iterates on guardrails before making Argon available to developers, enterprises, and consumers — starting with paid API customers and Google AI Ultra subscribers.
Internal results at Google
Argon is already powering internal Google workflows, with thousands of Googlers highlighting its strengths in specialized coding tasks, deeper research, and writing quality. Google detailed three concrete engineering results:
Quantum algorithmic optimization. Argon is helping Google's quantum computing researchers optimize the spacetime resources — qubits × gates — of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Memory efficiency. A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google's data centers. Google says the changes freed up over 300 TiB of memory once rolled out, with estimated total savings of 500 TiB to 1 PiB.
Large-scale codebase migrations. Argon agents are migrating C/C++ codebases to Rust across Google, scaling from tens of thousands of lines in core libraries like re2 and libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel. Given the criticality of these systems, Google says the rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before reaching production.
The libgav1 case is the most detailed. Argon agents took an existing Rust port of Google's open source video decoder and replaced 32K lines of SIMD code through many rounds of profile-guided experiments, studying the compiler's output and producing safe Rust that the compiler would vectorize automatically. The result: a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++ version.
Benchmarks
Google claims state-of-the-art results across coding, enterprise knowledge work, and video understanding:
- DeepSWE v1.1: 77.9% — a new state of the art on the benchmark measuring performance in real-world long-horizon software engineering tasks.
- Vals Index — Argon is the leading model on this index measuring economic impact across finance, coding, legal, and tax work, with each sector weighted by its contribution to U.S. GDP. Google reports similarly leading performance on Vals Finance Agent v2 (multi-step financial research) and Harvey's Legal Agent Benchmark (legal research and drafting).
- AutomationBench: 51.3%, rank #1 — Zapier's benchmark measuring end-to-end execution across core business functions.
- LVBench: 91.7% — state of the art on long video understanding.
Google also positions Argon as uniquely strong when knowledge work requires visual understanding, including professional chart analysis, identifying details from long videos, and taking action based on a series of documents.
Cybersecurity defense — without guardrails for defenders
The cybersecurity capabilities are the most consequential part of the announcement. Google trained Argon to autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and Google's internal teams, the company will release Argon without cyber guardrails.
Wiz is already using Argon through its Scan for Good initiative, a program that protects critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration, Google says the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — "a severe risk that previous frontier models had missed."
On CWE-bench v1, which evaluates a model's ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber's frontier performance on CWE-bench v0.
Google reports two additional internal results on vulnerability discovery. On Google's internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages. On Wiz's internal black-box penetration testing benchmark, which tests a model's ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
Safety work before broad release
Google is strengthening frontier safeguards across four areas before rolling Argon out broadly.
Defending against misuse. To prevent bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the model refuses harmful requests while preserving legitimate dual-use scientific research, per Google's Frontier Safety Framework. Google says it is improving techniques to monitor the model's internal activations to spot misuse, and that the safeguards underwent robustness testing by internal and external red teams using manual and automated attack methods.
Prompt injection defense. Google calls Argon its most resilient model yet against indirect prompt injections, in which malicious instructions or context hijack a model's behavior. Through automated red teaming and adversarial training, the company says Argon leads in prompt injection robustness on Gray Swan's Indirect Prompt Injection (IPI) benchmark.
Monitoring for misalignment. Google is deploying mitigations that monitor Argon's chain-of-thought and actions and stop execution when the model attempts to accomplish a task in ways that exceed the user's intentions. The company used a similar system to monitor its training runs, sending alerts to a dedicated incident response team while taking precautions against feeding findings back into training — a step taken, Google says, to avoid shaping Argon's reasoning to evade monitoring. The company also issued a direct call to the field: "We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment."
Hardening systems. In line with its agent control roadmap, Google is hardening sandboxed environments by isolating and sealing them before high-risk training or evaluations begin, and says it will share these agent security practices with partners.
What comes next
Google frames Argon as a partner for developers, professionals, and enterprises tackling the most difficult problems, with frontier-level capabilities in coding, knowledge work, cybersecurity defense, and creative writing. The immediate next step is the real-world evaluation cycle: the initial cohort of cyber defenders and trusted testers will stress the model and its guardrails before the release widens to paid API customers and Google AI Ultra subscribers. How Argon's unguarded cybersecurity capabilities perform in defenders' hands — and how the voluntary pre-release process with the U.S. government shapes that timeline — will define the rollout.
Original: vals.ai
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
170 articles
Related articles
- Google Announces Gemini 4 Argon, But No One Outside Can Use It
- Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
- Google Launches Gemini 3 Pro at $2/Million Input Tokens
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
- Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score