Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
Google's Gemini 3.1 Flash TTS ships today in preview with a 1,211 Elo score on Artificial Analysis, audio tags for mid-sentence control, 70+ languages, and SynthID watermarking.

Updated
Why it matters
- Gemini 3.1 Flash TTS achieved an Elo score of 1,211 on the Artificial Analysis TTS leaderboard, based on thousands of blind human preferences.
- The model rolls out in preview via the Gemini API and Google AI Studio for developers, Vertex AI for enterprises, and Google Vids for Workspace users.
- All generated audio is watermarked with SynthID for detection of AI-generated content.
Google is releasing Gemini 3.1 Flash TTS today, a text-to-speech model the company calls its most natural and expressive to date, with an Elo score of 1,211 on the Artificial Analysis TTS leaderboard — a benchmark built on thousands of blind human preferences.
The rollout covers three surfaces at launch. Developers get the model in preview through the Gemini API and Google AI Studio. Enterprises can access it in preview on Vertex AI. Workspace users receive it through Google Vids.
Artificial Analysis has placed Gemini 3.1 Flash TTS in its "most attractive quadrant," a position the benchmark operator reserves for models combining high-quality speech generation with low cost. That placement matters for buyers: TTS is becoming a cost-driven commodity market, and per-request economics increasingly decide which model wins enterprise voice workloads.
What the model does
The model supports native multi-speaker dialogue, more than 70 languages, and control over vocal style through natural language commands. Google says these optimizations bring style, pacing and accent control to major markets, letting developers build localized speech experiences at global scale.
The headline feature is a new system of audio tags. Developers embed natural language commands directly into text input to steer vocal style, pace and delivery. According to Google, this gives a level of granularity that earlier control schemes lacked.
Google frames the developer experience through a "director's chair" metaphor, with three configurable layers:
- Scene direction. Developers define the environment and provide specific dialogue instructions. Google says this world-building context helps characters remain "in-character" and react to one another naturally across multiple turns.
- Speaker-level specificity. Developers cast characters using unique Audio Profiles, then apply Director's Notes to toggle pace, tone and accent. Inline tags let speakers pivot from those high-level settings and change expression mid-sentence.
- Seamless export. Once a performance is tuned, developers export the exact parameters as Gemini API code, ensuring consistent, recognizable voices across projects and platforms.
Google positions these configurations as tools for precision in specific scenarios — creating memorable characters and immersive audio experiences.
Why it matters
Expressive, controllable TTS sits at the center of several fast-growing markets: AI agents with voice interfaces, audiobook and podcast automation, dubbing and localization, and interactive entertainment. Multi-speaker dialogue support and mid-sentence expression changes push the model toward scripted-content use cases that earlier TTS systems handled poorly.
The 70+ language count signals where Google expects volume. Localization has historically been one of the most expensive parts of global audio production, and cheap, expressive multi-language TTS compresses that cost.
Early developer and enterprise testers are already reporting results, according to Google. The company says testers have highlighted the model's controllability and expressivity, and described how audio tags provide a new level of creative precision — "transforming simple text into a high-fidelity vocal performance."
Provenance and safety
All audio generated by Gemini 3.1 Flash TTS is watermarked with SynthID. The watermark is imperceptible and interwoven directly into the audio output, allowing reliable detection of AI-generated content. Google frames this as a misinformation-prevention measure. The company has published a model card documenting its safety and responsibility approach.
The watermarking question carries growing policy weight. Regulators and platforms are converging on provenance requirements for synthetic media, and audio remains a common vector for voice-cloning fraud and deception. A detectable watermark baked into every output gives Google an answer it can point to — though detection only works where downstream platforms and tools actually check for it.
What to watch
The immediate question is competitive positioning. Google is claiming a leaderboard result and a price-quality sweet spot on the same day the model ships in preview, a pattern that mirrors how quickly TTS vendors now trade benchmark leads. Whether 3.1 Flash TTS holds the "most attractive quadrant" position will depend on how rivals respond on both quality and cost.
The second question is adoption depth. Preview access on Vertex AI means enterprise voice workloads can move quickly, and the export-as-API-code workflow is designed to lock tuned voice configurations into Google's stack. Developers can start experimenting with the audio tags and configurable controls now in the Google AI Studio Playground.
Original: aistudio.google.com
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
114 articles
Related articles
- Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
- Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
- Google's Flash TTS models build AI voices from text descriptions
- Google Ships Gemini 3.8 Live and Extended Thinking Models
- Google Ships Gemini 3.1 Flash Live Audio Model Globally