Models

Google Launches Gemini 3.1 Flash-Lite Starting at $0.25 Per Million Tokens

Google launches Gemini 3.1 Flash-Lite at $0.25/1M input tokens, with 2.5X faster first answers than 2.5 Flash and an Elo of 1432 on Arena.ai.

Gemini 3.1 Flash-Lite: Built for intelligence at scale
Gemini 3.1 Flash-Lite: Built for intelligence at scaleschoschie / Openverse
By Sophie Lindqvist2 min read

Updated

Why it matters

  • Gemini 3.1 Flash-Lite is priced at $0.25/1M input tokens and $1.50/1M output tokens, available in preview via the Gemini API in AI Studio and Vertex AI.
  • According to the Artificial Analysis benchmark, it delivers a 2.5X faster Time to First Answer Token and 45% higher output speed than 2.5 Flash at similar or better quality.
  • The model scores 1432 Elo on Arena.ai, 86.9% on GPQA Diamond and 76.8% on MMMU Pro, surpassing prior-generation larger models like 2.5 Flash.

Google today introduced Gemini 3.1 Flash-Lite, its fastest and most cost-efficient Gemini 3 series model, rolling out in preview to developers via the Gemini API in Google AI Studio and to enterprises via Vertex AI.

The model costs $0.25 per 1 million input tokens and $1.50 per 1 million output tokens. Google positions it as the budget tier of the Gemini 3 family, built for high-volume developer workloads where cost per call determines viability.

Speed gains over 2.5 Flash

According to the Artificial Analysis benchmark, 3.1 Flash-Lite delivers a 2.5X faster Time to First Answer Token and a 45% increase in output speed compared to 2.5 Flash, while maintaining similar or better quality. Google argues that latency at this level is required for high-frequency workflows, making the model suitable for developers building responsive, real-time experiences.

The speed-and-quality tradeoff matters commercially. Translation pipelines, content moderation and other always-on services process millions of calls per day, and small per-token savings compound into significant infrastructure costs. Small models that retain reasoning ability let companies route complex workloads to cheap tiers instead of premium ones.

Benchmark results

Gemini 3.1 Flash-Lite achieves an Elo score of 1432 on the Arena.ai Leaderboard, Google reports. It outperforms other models in its tier across reasoning and multimodal understanding benchmarks, scoring 86.9% on GPQA Diamond and 76.8% on MMMU Pro.

Google says the model surpasses larger Gemini models from prior generations, including 2.5 Flash itself. GPQA Diamond tests graduate-level scientific reasoning; MMMU Pro measures multimodal understanding across text and images.

Configurable "thinking" levels

3.1 Flash-Lite ships standard with thinking levels in AI Studio and Vertex AI. Developers can select how much the model "thinks" for a given task, a control Google describes as critical for managing high-frequency workloads.

The model targets two distinct workload classes. For cost-priority tasks at scale, Google cites high-volume translation and content moderation. For deeper reasoning, it points to generating user interfaces and dashboards, creating simulations and following complex instructions.

Early adopters

Latitude, Cartwheel and Whering are already using 3.1 Flash-Lite in early access on AI Studio and Vertex AI. Early testers highlighted the model's efficiency and reasoning capabilities, saying it can handle complex inputs with the precision of a larger-tier model, plus follow instructions and maintain adherence.

The release extends Google's tiered Gemini 3 strategy: large flagship models for frontier tasks, compact variants for scale. With preview pricing already undercutting most competitors in the small-model segment, expect developers to pressure-test whether 3.1 Flash-Lite's reasoning quality holds at production volume before committing high-frequency pipelines to it.

Original: aistudio.google.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

115 articles

Related articles

  1. Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant
  2. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  3. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
  4. Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
  5. Google Ships Gemini 3.1 Flash Live Audio Model Globally

« Previous articleNext article »