Google Ships Gemini 3.1 Flash Live Audio Model Globally
Google's Gemini 3.1 Flash Live tops audio benchmarks, doubles conversation memory, and powers Search Live's expansion to over 200 countries, all SynthID-watermarked.

Updated
Why it matters
- Gemini 3.1 Flash Live scores 90.8% on ComplexFuncBench Audio and 36.1% on Scale AI's Audio MultiChallenge with thinking enabled
- Search Live expands to more than 200 countries and territories powered by the multilingual 3.1 Flash Live model
- All 3.1 Flash Live audio output is watermarked with SynthID for AI-content detection; Verizon, LiveKit, and The Home Depot gave positive early feedback
Google has released Gemini 3.1 Flash Live, which the company calls its "highest-quality audio and voice model yet," and made it available simultaneously to developers, enterprises, and consumers. The rollout spans the Gemini Live API in Google AI Studio (developer preview), Gemini Enterprise for Customer Experience, and consumer surfaces Gemini Live and Search Live.
The release matters because voice is becoming the primary interface for AI assistants and customer-service automation. Real-time audio models must handle interruptions, hesitations, and long conversations — failure modes that have limited voice agents in production. Google is claiming measurable progress on exactly those problems.
Benchmark results
On ComplexFuncBench Audio, a benchmark testing multi-step function calling under various constraints, Google says 3.1 Flash Live scores 90.8%, leading against the company's previous model. The company did not name external competitors on that benchmark.
On Scale AI's Audio MultiChallenge, the model scores 36.1% with "thinking" enabled. Google describes that benchmark as specifically testing "complex instruction following and long-horizon reasoning amidst the interruptions and hesitations typical of real-world audio." A 36.1% score leading the field also indicates how far real-world audio understanding remains from solved.
Better tonal understanding for enterprises
Google says 3.1 Flash Live has improved tonal understanding compared to 2.5 Flash Native Audio, the model it effectively replaces in Google's audio stack. In Gemini Enterprise for Customer Experience, the model is "even more effective at recognizing acoustic nuances like pitch and pace" than its predecessor, and "better at dynamically adjusting its response to users' expressions of frustration or confusion."
That last capability carries direct commercial weight: a customer-service agent that detects frustration and adapts could reduce escalation rates in deployed call-center automation. Google reports positive feedback from three named early users — Verizon, LiveKit, and The Home Depot — who it says highlighted the model's "improved, natural conversation" in their workflows. The company did not publish detailed case-study results from those companies.
Consumer features and the Search Live expansion
For everyday users, Google promises two concrete gains in Gemini Live: faster responses than the previous model, and the ability to "follow the thread of your conversation for twice as long, keeping your train of thought intact during longer brainstorms." Google did not specify the exact context-window length behind that doubling.
The model is also "inherently multilingual," and Google used that property to drive this week's global expansion of Search Live. Users in more than 200 countries and territories can now hold "real-time, multimodal conversations with Search in their preferred language." That positions Search Live directly against voice-first assistants and conversational search products from competitors, extending Google's core search franchise into real-time voice across most of its global user base.
Google's blog post demonstrates the consumer use case with real-time troubleshooting help in Search Live — pointing users toward technical support scenarios rather than simple factual queries.
Watermarking with SynthID
Every piece of audio generated by 3.1 Flash Live carries a SynthID watermark. Google describes it as "imperceptible" and "interwoven directly into the audio output, allowing the reliable detection of AI-generated content to help prevent misinformation." The company directs readers to the model card for details on its safety and responsibility approach.
The watermarking commitment lands as regulators and platforms push for provenance standards for synthetic media. Audio deepfakes have driven recent fraud and election-misinformation incidents, making built-in detectability a competitive and compliance asset rather than a technical footnote.
What to watch
3.1 Flash Live is available starting today across all four surfaces. The open questions are ones Google's announcement does not answer: how the 90.8% and 36.1% benchmark leads translate into production reliability at enterprise scale, whether SynthID detection holds up against real-world audio processing and compression, and how Verizon, LiveKit, and The Home Depot deploy the model beyond pilot feedback. The pace of the release — a numbered point update going out simultaneously to developers, enterprises, and more than 200 countries — signals that Google intends to compete on distribution speed in voice AI, not just benchmark scores.
Original: ai.google.dev
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
- Google Ships Gemini 3.8 Live and Extended Thinking Models
- Google Launches Gemini 3.5 Live Translate Across Products
- Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
- Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month