Models

Google launches Gemini 3 Flash at $0.50 per million input tokens

Google launched Gemini 3 Flash today at $0.50 per million input tokens. The third release in its Gemini 3 family matches Gemini 3 Pro on multimodal reasoning and runs 3x faster than its predecessor.

Gemini 3 Flash: frontier intelligence built for speed
Gemini 3 Flash: frontier intelligence built for speedAI-generated
By Sophie Lindqvist6 min read

Updated

Why it matters

  • Gemini 3 Flash launches today at $0.50 per million input tokens and $3 per million output tokens, with audio input priced at $1 per million tokens
  • Model scores 90.4% on GPQA Diamond, 33.7% on Humanity's Last Exam without tools, and 81.2% on MMMU Pro — comparable to Gemini 3 Pro
  • Gemini 3 Flash runs 3x faster than Gemini 2.5 Pro per Artificial Analysis benchmarking and uses 30% fewer tokens on average
  • On SWE-bench Verified, Gemini 3 Flash scored 78%, above both the 2.5 series and Gemini 3 Pro
  • Gemini 3 Flash replaces 2.5 Flash as the default model in the Gemini app worldwide; Google API has processed more than 1 trillion tokens per day since the Gemini 3 rollout began

Google launched Gemini 3 Flash today at $0.50 per million input tokens, the third release in its Gemini 3 family and a model the company claims matches Gemini 3 Pro on multimodal reasoning while running three times faster than its predecessor.

The model is rolling out globally across Google's developer, consumer and enterprise surfaces. Pricing sits at $0.50 per million input tokens and $3 per million output tokens, with audio input at $1 per million tokens.

Google framed the launch as a step toward "frontier intelligence built for speed at a fraction of the cost." That phrase anchors the announcement and reflects how Google is positioning Flash against higher-priced reasoning models.

What does Gemini 3 Flash add to the lineup?

Gemini 3 Flash is the third member of the Gemini 3 family, following Gemini 3 Pro and Gemini 3 Deep Think, both released last month. Google built the Flash variant to keep "Pro-grade reasoning" while preserving Flash-level speed and cost.

The Flash designation dates back to Google's earliest Gemini releases and has traditionally prioritized latency and unit economics over frontier reasoning. With Gemini 3 Flash, Google is collapsing that gap. The new model's benchmark parity with Gemini 3 Pro on MMMU Pro is the clearest sign of that shift.

How does it score on the hard benchmarks?

Google disclosed three headline benchmark scores:

  • GPQA Diamond: 90.4% on a PhD-level reasoning and knowledge test that probes graduate-level science questions
  • Humanity's Last Exam: 33.7% without tools on a broad-coverage academic benchmark
  • MMMU Pro: 81.2% on a multimodal understanding test covering college-level problems across disciplines

Google called the MMMU Pro score "comparable to Gemini 3 Pro." That parity with the company's flagship on a multimodal benchmark is the central technical claim of the launch.

How fast and efficient is it?

Google said Gemini 3 Flash is "3x faster" than Gemini 2.5 Pro based on Artificial Analysis benchmarking. The model also uses 30% fewer tokens on average than 2.5 Pro when measured on typical traffic.

The model can "modulate how much it thinks," Google said. At the highest thinking level, it spends more compute on complex queries and pulls back on routine ones. Google plotted the performance gain as a shift on the Pareto frontier of quality versus cost, with the LMArena Elo Score as the quality axis. The claim: developers can reach frontier-quality outputs at Flash-tier prices, without paying for Pro-tier inference on every call.

Where can developers use it?

Google is shipping Gemini 3 Flash across six developer surfaces:

  • Gemini API in Google AI Studio for direct model access
  • Google Antigravity, the company's new agentic development platform launched alongside this release
  • Gemini CLI for command-line workflows
  • Android Studio for mobile app development
  • Vertex AI for enterprise machine-learning pipelines
  • Gemini Enterprise for business customers

The model is available in preview on developer surfaces today. On SWE-bench Verified, a coding-agent benchmark that scores models on real GitHub issues, Gemini 3 Flash scored 78% — above both the older 2.5 series and Gemini 3 Pro, Google said.

That result is unusual. Flash models typically trail their Pro counterparts on agentic coding tasks. A Flash score above Pro signals that Google's cost engineering is closing the gap between tiers.

Who is already deploying it?

Google named three early enterprise customers: JetBrains, Bridgewater Associates and Figma. The company did not detail specific deployments, but said each firm recognized the model's inference speed and reasoning capability.

The early adopter mix hints at Google's target markets:

  • JetBrains signals a play for coding-agent workloads and IDE integrations
  • Bridgewater points to financial-services deployments, likely research and decision-support tooling
  • Figma suggests creative-tooling and design application integrations

What changes for everyday consumers?

For everyday users, the most visible shift is the default-model swap. Gemini 3 Flash now powers the Gemini app worldwide, replacing 2.5 Flash at no additional cost.

Google pitched two consumer use cases for the upgrade:

  • Multimodal understanding. Users can upload videos and images and receive "a helpful and actionable plan in just a few seconds," Google wrote.
  • Voice-driven app building. Dictation can turn "unstructured thoughts into a functioning app in minutes," Google said, without requiring prior coding knowledge.

In Search, AI Mode is moving to Gemini 3 Flash as the default worldwide. The new model is "more powerful at parsing the nuances of your question," Google wrote. It serves "thoughtful, comprehensive responses" that combine research with action, the company said, targeting complex queries like last-minute trip planning or fast concept learning.

How big is the API business behind this launch?

Google said its API has processed more than 1 trillion tokens per day since the Gemini 3 rollout began. That figure is the largest daily token volume Google has disclosed for the family.

The volume signals the scale Google is operating at with Gemini 3. With Flash priced at $0.50 per million input tokens, the company can now capture high-volume, low-latency workloads that Pro pricing alone would not serve. The combination of speed, price, and benchmark parity positions Flash as the workhorse tier for production agentic systems.

What are the limits of the disclosed numbers?

Several caveats apply to the launch claims:

  • The 90.4% GPQA Diamond score came without a direct same-tier competitor comparison
  • The Humanity's Last Exam figure — 33.7% without tools — leaves tool-augmented results unpublished
  • The "3x faster" claim rests on Artificial Analysis benchmarking, whose methodology Google did not restate
  • Token-efficiency claims anchor to "typical traffic," not a standardized workload
  • Google did not disclose context window, absolute latency, or safety-evaluation specifics

None of these omissions undercut the headline numbers, but they leave room for independent verification and for competitors to publish contrasting benchmarks.

What comes next?

The full Gemini 3 family — Pro, Deep Think and now Flash — is now live across consumer, developer and enterprise surfaces. Google closed the announcement with a direct invitation: "We're looking forward to seeing what you bring to life with this expanded family of models."

That closing line doubles as a market signal. With Flash in market, Google now covers the full reasoning cost curve from low-latency agentic workflows to the hardest multi-step problems. The trillion-token daily volume on Gemini 3 suggests the company intends to make every tier of that curve a meaningful revenue line. Watch for third-party benchmark evaluations and pricing moves from competing models in the coming weeks.

Original: blog.google

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

209 articles

Related articles

  1. Google Cuts Gemini Flash Pricing in Half With 3.7 Release
  2. Google Ships Gemini 3.6 Flash and 3.5 Flash-Lite, Targets Agent Costs
  3. Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens
  4. Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
  5. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month

« Previous article