Safety & Security

Google Releases Gemini 3.5 Flash Cyber for Vulnerability Hunting

Google's new Gemini 3.5 Flash Cyber finds 55 unique V8 bugs versus Opus 4.6's 36, and ships only to governments via CodeMender under a limited-access pilot.

Introducing Gemini 3.5 Flash Cyber
Introducing Gemini 3.5 Flash CyberAI-generated
By Sophie Lindqvist4 min read

Updated

Why it matters

  • Gemini 3.5 Flash Cyber is a cybersecurity model fine-tuned on 3.5 Flash, available only to governments and trusted partners via CodeMender under a limited-access pilot program.
  • On the V8 JavaScript Engine, 3.5 Flash Cyber found 55 unique confirmed issues, versus 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6, including 10 issues the other models missed.
  • Google's Cloud Vulnerability Research team used the model to find remote code execution flaws and a memory-corruption bug in 2 hours, including a 100% reliable RCE exploit bypassing ASLR and W^X.

Google has introduced Gemini 3.5 Flash Cyber, a lightweight cybersecurity model fine-tuned on top of 3.5 Flash to find, validate, and patch software vulnerabilities faster and more cheaply than its mainline Flash models. Access will be restricted: the model ships only to governments and trusted partners through Google's CodeMender agent as part of a limited-access pilot program, a deliberate constraint the company says addresses the dual-use risk of a model that can weaponize the flaws it finds.

The stakes are straightforward. Google argues that AI agents are now finding vulnerabilities faster than human defenders can fix them, and that closing that gap requires security models that are cheap and scalable enough to run constantly, not just occasionally. The company positions 3.5 Flash Cyber as a cost-efficient alternative to the large, expensive cybersecurity models competitors offer.

The engineering logic behind the model is the core of the announcement. Finding deep flaws in large codebases means exploring an immense execution search space, and a single expensive call to a massive model becomes a bottleneck. Because Flash Cyber is fast and cheap, CodeMender invokes it multiple times, up to five calls per final report, letting sub-agents analyze far more code paths and merge their findings into a single high-quality report. That architecture, Google says, makes the model easy to embed in frequent scans, time-sensitive launch processes, and commit-scanning pipelines at scale.

Google tested the model on three benchmarks. On CyberGym, which evaluates agents against hundreds of real-world software vulnerabilities, the multi-invocation CodeMender agent built on 3.5 Flash Cyber achieved competitive performance against significantly larger models. Google notes that competitor results on CyberGym are self-reported scores.

On a separate evaluation built independently by Google's Big Sleep team, focused on critical, hard-to-find vulnerabilities in complex codebases like Chrome and Safari, 3.5 Flash Cyber significantly surpassed both mainline 3.5 Flash and 3.6 Flash. Google says this stress test ran without safety guardrails.

The third test came from Google Chrome's production commit scanning pipeline. Because the vulnerabilities tested there were never publicly disclosed, Google says the benchmark remained free of contamination for both Gemini and competitor models. 3.5 Flash Cyber showed a significant uplift over 3.5 Flash. Google also reports that more recent competitor model versions after Opus 4.6 refuse to fulfill these tasks due to built-in safety guardrails, and are therefore not shown in the results.

The unique-discovery numbers illustrate why Google favors a wider net. When tested on the V8 JavaScript Engine across a fixed number of invocations, 3.5 Flash Cyber found 55 unique confirmed issues, versus 47 for mainline 3.5 Flash and 36 for Claude Opus 4.6, including 10 issues neither of the other two models caught. Google's explanation: weaker models get stuck in loops, repeatedly finding the same issue while missing critical ones, while a strong cheap model keeps discovering new code paths as invocations scale.

The model is already working inside Google. CodeMender running 3.5 Flash Cyber is finding and fixing vulnerabilities in Google's internal codebases, including Chrome, Android, Cloud, Ads, and YouTube. Google's Cloud Vulnerability Research team used the model to proactively secure systems in record time: in just 2 hours, it uncovered remote code execution vulnerabilities in public APIs and found a memory-corruption vulnerability in a sensitive production service. It then generated what Google describes as a 100% reliable remote-code execution exploit that bypassed standard mitigations including Address Space Layout Randomization (ASLR) and Write XOR Execute (W^X). Early feedback from testers at Wiz and Cloud CISO Security Engineering confirms a significant capability improvement over mainline 3.5 Flash, according to Google.

Google attributes the model's training quality to its own security data assets: OSV.dev, a vulnerability database it runs spanning over 700,000 open-source vulnerabilities, and more than ten years of OSS-Fuzz results. That data lets the company train on real security work rather than synthetic examples, teaching models to operate industry-standard tools, read millions of lines of code in projects like Chromium, and sustain hours of continuous, deep analysis.

Google is also bringing CodeMender's foundational capabilities to customers more broadly through the Gemini Enterprise Agent Platform with generally available Gemini models. That split, a locked-down specialist model for states and vetted partners, and general agent tooling for enterprises, defines how Google plans to scale AI-driven defense without handing offensive capability to everyone.

Original: cloud.google.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

115 articles

Related articles

  1. Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
  2. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
  3. Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens
  4. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  5. Google Confirms Gemini Hacked Three Real Companies in May Test

Next article »