Google Cuts Gemini Flash Pricing in Half With 3.7 Release
Google released Gemini 3.7 Flash at half the price of its predecessor, posting double-digit gains on five benchmarks including DeepSWE v1.1 and GDP.pdf, while updating Spark and its CBRN safety safeguards.

Updated
Why it matters
- Gemini 3.7 Flash released three weeks after Gemini 3.6 Flash
- Introductory price: $0.75 per million input tokens and $3.75 per million output tokens through end of 2025, half the 3.6 Flash rate
- DeepSWE v1.1 score rose to 65.3% from 49.0% on 3.6 Flash
- FrontierCode 1.1 Main climbed to 43.6% from 34.4%
- Gemini Spark, available in 160+ countries to Google AI Pro and Ultra subscribers, switches to 3.7 Flash on day one
Google released Gemini 3.7 Flash on Tuesday, pricing the model at half the per-token cost of its predecessor while posting double-digit gains on coding and reasoning benchmarks.
The new model enters general availability three weeks after Gemini 3.6 Flash, a release cadence that reflects pressure from developers asking for rapid iteration, according to Google. "Our most intelligent workhorse model yet for coding and agents," the company wrote in its release notes, framing the upgrade as a direct response to user feedback and what it called "algorithmic innovations that we look forward to bringing to future models."
The pricing math is the headline number. Gemini 3.7 Flash lists at $0.75 per million input tokens and $3.75 per million output tokens through the end of the year — exactly half the introductory rate Google charged for 3.6 Flash. The company has not said what the steady-state price will be once the promotion ends in January.
What did the benchmarks actually measure?
Google anchored its performance claims in five external evals spanning coding, web development, and knowledge work.
On FrontierCode 1.1 Main, which measures end-to-end software engineering tasks such as debugging and issue resolution, 3.7 Flash climbed to 43.6% from 34.4% on 3.6 Flash. On DeepSWE v1.1, a software engineering benchmark from the same family, the model scored 65.3% versus 49.0%. The roughly 16-point jump on DeepSWE is the largest single delta Google disclosed.
Web development is where the gains look most commercially relevant. 3.7 Flash posted an Elo of 1588 on Arena.ai's WebDev Arena, up 50 points from 1538 on 3.6 Flash. Google also reports the model produces more functional, feature-complete layouts in fewer prompts and tracks reference designs — screenshots, images, or full design systems — more faithfully than its predecessor.
For knowledge-dense fields, the model improved on the GDP.pdf benchmark, which tests a model's ability to process long, complex documents, to 34.0% from 22.0%. On AutomationBench, a measure of real-world business workflow completion, it reached 30.4% versus 17.0% — nearly double.
Why does the price cut matter?
The combination of higher accuracy and lower per-token cost is the calculation Google is asking developers to make. Agents that previously cost more to run per query can now complete longer task chains at lower cost. The pitch is that lower unit economics — not raw capability alone — will decide which models carry enterprise traffic.
The Flash line has historically been Google's mid-tier workhorse, sitting below Pro and above the discontinued Nano tier. It carries the bulk of Gemini's API volume because it can be deployed cheaply in high-traffic contexts. Cutting its price in half while lifting accuracy on agentic benchmarks is a direct lever on production cost-per-task.
Google also points to qualitative changes. The model "better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity," the company wrote. It "thinks more diligently, putting in more effort into multi-step planning and tool calls." Those claims are harder to verify from external benchmarks, but they are the changes developers care about most when chaining agents.
How does this fit Spark?
Gemini Spark, the personal assistant Google launched at I/O in May, will run on 3.7 Flash from Tuesday onward. The agent is available to Google AI Pro and Ultra subscribers in more than 160 countries and operates as "your personal AI agent that runs 24/7, taking action on your behalf while under your direction," according to Google's product page.
The model upgrade targets Spark's weakest link: reliable tool use across Workspace apps. Google says 3.7 Flash improves tool calling for Gmail, Drive, and Docs, which matters because Spark's headline use cases — consolidating files, drafting emails, updating status documents — depend on chaining those tools in sequence without supervision.
Spark is also the consumer face of Google's agent strategy. Most users will never touch the Gemini API. They will judge Google's agent ambitions by whether Spark drafts a coherent email, updates a Sheet, and follows up on a calendar invite without dropping a step.
What about safety?
The release ships with updated safeguards in two risk domains Google has flagged as priorities: chemical, biological, radiological, and nuclear misuse, and cyber offense. The company's Frontier Safety framework was extended to cover both areas. A new model card documents the mitigations, though Google has not disclosed which specific red-team findings drove each change — a recurring gap in how frontier model providers document safety updates.
The cyber and CBRN categories are the two areas Google has committed to monitoring publicly. They are also the categories where open-weight competitors have drawn the most scrutiny. By tying the 3.7 Flash release to explicit safeguards, Google is signaling it intends to ship agentic-capable models without reopening the debate over capability overhang.
Where can developers get it?
Google is routing 3.7 Flash through three channels.
Developers can test it in Google Antigravity, the agent-first coding surface, or build against it through the Gemini API in Google AI Studio and Android Studio. Enterprises reach the model through Gemini Enterprise Agent Platform and the Gemini Enterprise app. Consumers touch it through Spark in the Gemini app, where it serves paying Pro and Ultra subscribers.
The release lands in a market where Anthropic, OpenAI, and Meta are also pushing faster, cheaper mid-tier models. Claude Haiku, GPT-5 mini, and recent Llama variants compete in roughly the same price band. Google's bet is that iteration speed — three weeks between 3.6 and 3.7 — becomes a defensible advantage when the underlying model is good enough for production traffic.
The promotion runs through the end of December. What happens to pricing in 2026 will determine whether the Flash line stays at parity with rivals or reverts to a premium tier, and whether the agentic workloads it unlocks remain economically viable at the new steady-state rate.
Original: blog.google
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
207 articles
Related articles
- Google Ships Gemini 3.6 Flash and 3.5 Flash-Lite, Targets Agent Costs
- Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
- Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens
- Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
- Google Launches Gemini 3 Pro at $2/Million Input Tokens