Google Ships Gemini 3.6 Flash and 3.5 Flash-Lite, Targets Agent Costs
Gemini 3.6 Flash cuts output tokens 17% versus 3.5 Flash at lower prices, while Flash-Lite hits 350 tokens/s and a cyber model stays restricted to governments.

Updated
Why it matters
- Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash (Artificial Analysis Index), up to 65% on DeepSWE, priced at $1.50/$7.50 per 1M input/output tokens.
- 3.5 Flash-Lite runs at 350 output tokens/s, costs $0.30/$2.50 per 1M tokens, and outperforms 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%).
- 3.5 Flash Cyber is available only to governments and trusted partners via CodeMender in a limited-access pilot; Google has started pre-training for Gemini 4.
Google has released Gemini 3.6 Flash, a workhorse model that cuts output token usage by 17% compared to 3.5 Flash on the Artificial Analysis Index — while charging less per output token. Alongside it, the company launched 3.5 Flash-Lite, its fastest 3.5-class model at 350 output tokens per second, and 3.5 Flash Cyber, a security-specialized model available only to governments and trusted partners through a limited-access pilot.
The three releases target one pressure point in the AI market: developers running production agents need lower cost per task and lower latency, not just higher benchmark scores. Google positions its Flash series as "the sweet spot of efficiency and quality to enable scaling agentic workflows."
3.6 Flash: cheaper, leaner, better
3.6 Flash builds on developer feedback from 3.5 Flash. It takes fewer reasoning steps and fewer tool calls to complete multi-step workflows, and Google pairs that efficiency with a price cut: $1.50 per 1M input tokens and $7.50 per 1M output tokens, reducing the overall cost per agentic task.
The efficiency gains are steep on some benchmarks. Google reports up to 65% fewer output tokens on DeepSWE by Datacurve. Quality also improves across use cases:
- Coding and research: DeepSWE precision rises to 49% from 37%, with fewer unwanted code edits and reduced execution loops. MLE Bench improves to 63.9% from 49.7%.
- Computer use: OSWorld-Verified climbs to 83.0% from 78.4%. Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise.
- Knowledge work: GDPval-AA v2 rises to 1421 from 1349. Google names Hebbia and Harvey as customers using the model for document parsing, chart and data analysis, and report drafting.
The model ships with enhanced Frontier Safety safeguards covering chemical, biological, radiological, and nuclear (CBRN) risks and cyber offense misuse. Google says the safeguards make the model "substantially more resistant to jailbreaks" while training minimizes refusals for beneficial uses. A model card is available.
3.5 Flash-Lite: throughput at $0.30 per million tokens
3.5 Flash-Lite targets low-latency and high-throughput workloads such as agentic search and document processing. At $0.30 per 1M input tokens and $2.50 per 1M output tokens, it runs at 350 output tokens per second as measured by Artificial Analysis — the fastest model in the 3.5 series.
The gains over 3.1 Flash-Lite are large: Terminal-Bench 2.1 jumps to 54% from 31%, long-context benchmark GDM-MRCR v2 to 72.2% from 60.1%, and GDPval-AA v2 to 1140 from 642.
The model even beats the larger 3 Flash on several agentic and coding evals, including SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%). Google pitches it as a faster, more capable option for workloads currently on 2.5 and 3 Flash.
Developers can configure thinking levels per workload — minimal or low for cheap, high-volume execution, higher levels for multi-step subagent workloads. Computer use is also a built-in tool.
3.5 Flash Cyber: frontier security, restricted access
Google frames the cyber model around a gap it says is widening: "AI models have become capable of finding security vulnerabilities faster than current systems can fix them."
Built on 3.5 Flash and fine-tuned for finding and fixing vulnerabilities, 3.5 Flash Cyber runs inside CodeMender, Google's code security agent, which coordinates multiple model instances into a single combined report. Google says the combination reaches "competitive performance at the frontier" on the CyberGym benchmark, at a lower price per token than larger models.
Because the technology is dual-use, access is deliberately narrow. The model will be available exclusively to governments and trusted partners via CodeMender, as part of a limited-access pilot. Google says the approach "will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse."
Availability
3.6 Flash and 3.5 Flash-Lite are available starting today:
- Developers: via the Gemini API in Google AI Studio and Android Studio; 3.6 Flash also in Google Antigravity.
- Enterprises: in the Gemini Enterprise Agent Platform; 3.6 Flash also in the Gemini Enterprise app.
- Consumers: via the Gemini app; 3.5 Flash-Lite is also rolling out in Google Search.
Google also confirmed that Gemini 3.5 Pro is currently testing with partners and will be broadly available "as soon as it's ready." The bigger signal for the competitive race with OpenAI and Anthropic sits further out: Google says it has "started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress."
With 3.6 Flash, the efficiency story is now measurable in dollars per agent task — and the Gemini 4 pre-training run sets the timeline for Google's next frontier push.
Original: artificialanalysis.ai
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
117 articles
Related articles
- Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens
- Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
- Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
- Google Launches Gemini 3 Pro at $2/Million Input Tokens
- Google Releases Gemini 3.5 Flash Cyber for Vulnerability Hunting