Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
Google's updated Gemini 2.5 Flash Native Audio leads ComplexFuncBench Audio at 71.5%, and a live speech translation beta lands in Google Translate with 70+ languages.

Updated
Why it matters
- Gemini 2.5 Flash Native Audio scores 71.5% on ComplexFuncBench Audio, leading prior versions and competitors on multi-step function calling.
- Instruction adherence rose from 84% to 90%, and the model is generally available on Vertex AI and in preview in the Gemini API.
- A live speech-to-speech translation beta covering 70+ languages and 2,000 language pairs starts rolling out today in the Google Translate app on Android in the US, Mexico and India, with the Gemini API to follow in 2026.
Google has released an updated Gemini 2.5 Flash Native Audio model for live voice agents, scoring 71.5% on ComplexFuncBench Audio — a benchmark the company says puts it ahead of previous versions and industry competitors on multi-step function calling.
The release follows an upgrade earlier this week to the Gemini 2.5 Pro and Flash text-to-speech models. Google framed the two announcements as complementary: expressive speech generation on one side, and the conversational model that responds on the other. The updated native audio model is available now across Google AI Studio and Vertex AI, generally available on Vertex AI and in preview in the Gemini API. It has also started rolling out in Gemini Live and, for the first time, Search Live.
The stakes are commercial. Voice agents are emerging as one of the largest enterprise applications of generative AI, and Google is competing with OpenAI, Amazon and a crowded field of startups to become the default speech model underneath customer service, sales and support workloads.
Three upgrades to the voice model
Google says it improved Gemini 2.5 Native Audio in three areas.
Sharper function calling. The model now more reliably identifies when to fetch real-time information mid-conversation and weaves that data back into the audio response without breaking the flow. On ComplexFuncBench Audio, an evaluation that captures multi-step function calling under various constraints, Gemini 2.5 Native Audio leads with a score of 71.5%.
Robust instruction following. The model now achieves a 90% adherence rate to developer instructions, up from 84%, which Google says produces higher user satisfaction on content completeness.
Smoother conversations. The model retrieves context from previous turns more effectively, creating more cohesive multi-turn conversations.
Customer deployments
Google Cloud customers are already running the model in production, from mortgage processing to customer calls.
Shopify's Sidekick voice assistant loses the AI tell quickly, according to David Wurtz, VP of Product at Shopify: "Users often forget they're talking to AI within a minute of using Sidekick, and in some cases have thanked the bot after a long chat…New Live API AI capabilities offered through Gemini [2.5 Flash Native Audio] empower our merchants to win."
At United Wholesale Mortgage, the model powers Mia, a loan-generation agent launched in May 2025. "By integrating the Gemini 2.5 Flash Native Audio model…we've significantly enhanced Mia's capabilities since launching in May 2025. This powerful combination has enabled us to generate over 14,000 loans for our broker partners," said Jason Bressler, Chief Technology Officer at UWM.
Newo.ai uses the model through Vertex AI for its AI Receptionists. "They can identify the main speaker even in noisy settings, switch languages mid-conversation, and sound remarkably natural and emotionally expressive," said David Yang, Co-founder of Newo.ai.
Live speech translation in headphones
The second announcement targets consumers: live speech translation, a beta capability that streams speech-to-speech translation directly to headphones while preserving the speaker's intonation, pacing and pitch.
The feature supports two modes. Continuous listening automatically translates speech in multiple languages into a single target language — put headphones in and hear the surroundings in your own language. Two-way conversation mode translates between two languages in real time, switching the output language based on who is speaking. If an English speaker talks with a Hindi speaker, each hears their own language: English in the headphones, Hindi broadcast from the phone.
Google lists five capabilities for the translation system:
- Language coverage: translation across over 70 languages and 2,000 language pairs, combining Gemini's world knowledge and multilingual abilities with native audio.
- Style transfer: preservation of the speaker's intonation, pacing and pitch so translations sound natural.
- Multilingual input: understanding multiple languages simultaneously in a single session without switching settings.
- Auto detection: identifying the spoken language and starting translation without the user selecting one.
- Noise robustness: filtering ambient noise for conversations in loud, outdoor environments.
The beta is rolling out starting today in the Google Translate app. Users connect headphones and tap "Live translate." The experience is available on Android devices in the US, Mexico and India, with iOS support and more regions promised soon.
Translation is a market where Google has competed for two decades, and real-time headphone translation puts it in direct contention with similar efforts from other AI vendors chasing the hardware-adjacent use case.
What comes next
Google says it will iterate on the translation experience based on feedback and bring it to more Google products, including the Gemini API, in 2026. Developers can start building with Gemini 2.5 Flash Native Audio today in Google AI Studio, or via the Gemini API speech generation docs, prompting guide and Gemini API Cookbook. The Gemini 2.5 Flash and 2.5 Pro text-to-speech models remain available through the Gemini API in Google AI Studio.
The roadmap signals Google's bet that native audio — one model handling speech input and output directly, rather than stitching together transcription, text reasoning and synthesis — becomes the default architecture for both enterprise agents and consumer translation.
Original: blog.google
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
147 articles
Related articles
- Google Ships Gemini 3.8 Live and Extended Thinking Models
- Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
- Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
- Google Ships Gemini 3.1 Flash Live Audio Model Globally
- Google Launches Gemini 3.5 Live Translate Across Products