Google's Flash TTS models build AI voices from text descriptions
Google's Gemini 3.8 Flash TTS and Flash-Lite TTS cover 100+ languages, create voices from text descriptions, take stage directions, and clone voices from 30-second samples.

Updated
Why it matters
- Google introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, supporting more than 100 languages.
- Flash TTS can create new voices from text descriptions; both models support stage directions per line and two-voice dialogue from a single script.
- A voice cloning feature builds a voice profile from a 30-second audio sample.
Google has introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, that support more than 100 languages and can create entirely new voices from plain text descriptions. The release pushes synthetic voice technology closer to a design tool: instead of selecting from a preset catalog, developers describe the voice they want and the model generates it.
The headline capability sits in Gemini 3.8 Flash TTS. The model can construct a voice from scratch based on a written description, which moves voice creation from recording studios and lengthy samples to a prompt. Both new models also accept stage directions attached to individual lines of a script, giving fine-grained control over delivery within a single generation. That positions the models for scripted audio production rather than simple narration.
Dialogue is a first-class feature. Both Flash TTS and Flash-Lite TTS can generate two-voice conversations from a single script, removing the need to synthesize each speaker separately and stitch the results together. According to Google, a voice cloning feature complements the text-description workflow by building a voice profile from a 30-second audio sample.
The combination of features addresses the main bottlenecks in AI audio production: casting, direction, and multi-speaker assembly. A 30-second cloning threshold is short enough to make voice replication practical at scale, which is also where the technology's risks concentrate. Voice cloning has drawn regulatory scrutiny, and tools that lower the barrier to convincing synthetic speech tend to intensify debates over consent, impersonation, and disclosure.
The stakes for Google are competitive. Text-to-speech has become a battleground feature across the AI industry, powering assistants, audiobooks, dubbing, and accessibility tools. Multilingual coverage across more than 100 languages gives the Flash TTS models reach into markets where voice interfaces often outpace text input. The two-tier lineup, with Flash-Lite as the lighter option, suggests Google is targeting both latency-sensitive applications and higher-quality production workloads within the Gemini ecosystem.
The release signals where Google sees audio heading: toward prompt-driven generation, where a written description of a voice, a script with embedded directions, and a cloned sample are all interchangeable inputs to the same system. As these models roll out, the industry's attention will shift to how Google gates the cloning feature and whether the 30-second threshold comes with safeguards against misuse.
Original: docs.cloud.google.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles
Related articles
- Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
- Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
- Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
- OpenAI ships three realtime audio models, led by GPT-Realtime-2
- Google Rolls Out Gemini 3.8 Live with Live Avatar for Enterprise