Models

Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month

Google has released Gemini 3.5 Flash globally, claiming it beats Gemini 3.1 Pro on agentic benchmarks while running four times faster than frontier rivals. Gemini 3.5 Pro follows next month.

Gemini 3.5: frontier intelligence with action
Gemini 3.5: frontier intelligence with actionAI-generated
By Marcus Bennett5 min read

Updated

Why it matters

  • Gemini 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), MCP Atlas (83.6%), and CharXiv Reasoning (84.2%)
  • Google says 3.5 Flash is 4x faster in output tokens per second than other frontier models and often costs less than half as much for long-horizon tasks
  • Gemini 3.5 Pro is already used internally at Google and rolls out next month; Gemini Spark, a 24/7 personal agent built on 3.5 Flash, enters Beta for US Google AI Ultra subscribers next week

Google has released Gemini 3.5 Flash, the first model in its new Gemini 3.5 family, and made it available today to billions of users through the Gemini app and AI Mode in Google Search, alongside developer and enterprise channels.

The company positions the release as a turning point for AI agents rather than another incremental chatbot upgrade. "Today, we're introducing Gemini 3.5, our latest family of models combining frontier intelligence with action," Google said in its announcement. "This represents a major leap forward in building more capable, intelligent agents."

The stakes are straightforward. Agents that can plan, write code, and execute multi-step workflows without constant supervision are where Google, OpenAI, Anthropic, and others are now competing for developer and enterprise spending. With 3.5 Flash, Google is arguing that frontier-level capability and low latency no longer require a trade-off.

Benchmark results

Google claims 3.5 Flash outperforms Gemini 3.1 Pro — a larger model from the previous generation — on several demanding coding and agentic benchmarks:

  • Terminal-Bench 2.1: 76.2%
  • GDPval-AA: 1656 Elo
  • MCP Atlas: 83.6%
  • CharXiv Reasoning (multimodal understanding): 84.2%

The speed claim is the more unusual one. Measured in output tokens per second, Google says 3.5 Flash is four times faster than other frontier models. On the Artificial Analysis index, the model lands in the top-right quadrant — the region where high intelligence and high speed coincide — which Google reads as proof that "you no longer have to trade quality for latency."

Built for long-horizon work

The speed-performance balance matters most for long-horizon agentic tasks: jobs that take many steps, many tools, and many hours. "What used to take a developer days or an auditor weeks, 3.5 Flash can now help complete in a fraction of the time, often at less than half the cost of other frontier models," Google said. The company cites application development, codebase maintenance, and preparation of financial documents as concrete use cases.

Paired with an updated version of Google's Antigravity harness, 3.5 Flash can deploy collaborative subagents — multiple model instances working together on parts of a problem — under human supervision. Google says the setup can reliably execute multi-step workflows and coding tasks "while sustaining frontier performance." The model also builds on Gemini 3's multimodal foundation to generate richer, more interactive web UIs and graphics.

Google says it developed the 3.5 series in close collaboration with industry partners to locate where "toil and complexity" accumulated in real workflows. The payoff, according to the company, spans "banks and fintechs automating multi-week workflows to data science teams unearthing insights amidst complex data environments."

Availability

Gemini 3.5 Flash is generally available today across four surfaces:

  • Consumers: the Gemini app and AI Mode in Google Search, globally
  • Developers: Google Antigravity, the Gemini API in Google AI Studio, and Android Studio
  • Enterprises: Gemini Enterprise Agent Platform and Gemini Enterprise

The consumer rollout is effectively immediate: 3.5 Flash is now the default model behind the Gemini app and AI Mode in Search worldwide.

Gemini Spark, a personal agent

At its I/O event today, Google demonstrated how 3.5 Flash's agentic capabilities translate into consumer features. The headline product is Gemini Spark, described by Google as "your personal AI agent." It runs continuously. "It runs 24/7, helping you navigate your digital life, taking action on your behalf while under your direction," the company said.

Gemini Spark starts rolling out to trusted testers today. Google plans to bring a Beta to Google AI Ultra subscribers in the United States next week.

In Search, 3.5 Flash's agentic coding capabilities are powering new information agents that "work for you 24/7" and more dynamic generative UI experiences. Google demonstrated Search using the model to build an interactive visual explanation of Gyroid patterns.

What's next: 3.5 Pro

Google is not stopping at Flash. "We're also hard at work on 3.5 Pro. It's already being used internally, and we look forward to rolling it out next month," the company said. That internal-first deployment mirrors how Google handled earlier Gemini generations before public release, and it signals the company is already stress-testing the larger model on production workloads.

Safety posture

Gemini 3.5 was developed under Google's Frontier Safety Framework. The company says it has strengthened its cyber and CBRN (chemical, biological, radiological, nuclear) safeguards, making the model "less likely to generate harmful content, and to mistakenly refuse to answer safe queries."

The last point is a notable admission that over-refusal — models declining benign requests — has been a real cost of safety tuning. Google attributes the improvements to "new, more advanced safety training and mitigations, including interpretability tools that help check and understand the AI's inner reasoning before it provides a response." Interpretability-based mitigation, applied before a model responds, is an approach Google has increasingly leaned on to justify frontier releases under regulatory scrutiny.

Why it matters

The release condenses three competitive fronts into one launch. First, a small fast model beating its own larger predecessor on agentic benchmarks (3.5 Flash over Gemini 3.1 Pro) pressures rivals who still gate frontier capability behind premium-tier models. Second, sub-two-day availability across consumer, developer, and enterprise channels gives Google distribution few competitors can match. Third, Gemini Spark moves Google directly into the personal-agent category, where the model acts rather than merely answers.

With 3.5 Pro promised next month and already running internally at Google, the Flash release looks less like a standalone launch and more like the opening move of a coordinated rollout — one that will test whether frontier-speed agents, deployed at billions-of-users scale, can hold up under real workloads.

Original: blog.google

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

108 articles

Related articles

  1. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  2. Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
  3. Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens
  4. Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
  5. Google Recaps 2025: Gemini 3, AlphaFold Milestones and a Physics Nobel

« Previous articleNext article »