Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens
Google launches Gemini 3.1 Flash-Lite at $0.25/1M input tokens, with 2.5X faster first answers than 2.5 Flash and an Elo of 1432 on Arena.ai.

Updated
Why it matters
- Gemini 3.1 Flash-Lite is priced at $0.25/1M input tokens and $1.50/1M output tokens, available in preview via the Gemini API in AI Studio and Vertex AI.
- According to the Artificial Analysis benchmark, it delivers a 2.5X faster Time to First Answer Token and 45% higher output speed than 2.5 Flash at similar or better quality.
- The model scores 1432 Elo on Arena.ai, 86.9% on GPQA Diamond and 76.8% on MMMU Pro, surpassing prior-generation larger models like 2.5 Flash.
Google today introduced Gemini 3.1 Flash-Lite, its fastest and most cost-efficient Gemini 3 series model, rolling out in preview to developers via the Gemini API in Google AI Studio and to enterprises via Vertex AI.
The model costs $0.25 per 1 million input tokens and $1.50 per 1 million output tokens. Google positions it as the budget tier of the Gemini 3 family, built for high-volume developer workloads where cost per call determines viability.
Speed gains over 2.5 Flash
According to the Artificial Analysis benchmark, 3.1 Flash-Lite delivers a 2.5X faster Time to First Answer Token and a 45% increase in output speed compared to 2.5 Flash, while maintaining similar or better quality. Google argues that latency at this level is required for high-frequency workflows, making the model suitable for developers building responsive, real-time experiences.
The speed-and-quality tradeoff matters commercially. Translation pipelines, content moderation and other always-on services process millions of calls per day, and small per-token savings compound into significant infrastructure costs. Small models that retain reasoning ability let companies route complex workloads to cheap tiers instead of premium ones.
Benchmark results
Gemini 3.1 Flash-Lite achieves an Elo score of 1432 on the Arena.ai Leaderboard, Google reports. It outperforms other models in its tier across reasoning and multimodal understanding benchmarks, scoring 86.9% on GPQA Diamond and 76.8% on MMMU Pro.
Google says the model surpasses larger Gemini models from prior generations, including 2.5 Flash itself. GPQA Diamond tests graduate-level scientific reasoning; MMMU Pro measures multimodal understanding across text and images.
Configurable "thinking" levels
3.1 Flash-Lite ships standard with thinking levels in AI Studio and Vertex AI. Developers can select how much the model "thinks" for a given task, a control Google describes as critical for managing high-frequency workloads.
The model targets two distinct workload classes. For cost-priority tasks at scale, Google cites high-volume translation and content moderation. For deeper reasoning, it points to generating user interfaces and dashboards, creating simulations and following complex instructions.
Early adopters
Latitude, Cartwheel and Whering are already using 3.1 Flash-Lite in early access on AI Studio and Vertex AI. Early testers highlighted the model's efficiency and reasoning capabilities, saying it can handle complex inputs with the precision of a larger-tier model, plus follow instructions and maintain adherence.
The release extends Google's tiered Gemini 3 strategy: large flagship models for frontier tasks, compact variants for scale. With preview pricing already undercutting most competitors in the small-model segment, expect developers to pressure-test whether 3.1 Flash-Lite's reasoning quality holds at production volume before committing high-frequency pipelines to it.
Original: aistudio.google.com
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles