Research

Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction

Google launches Gemini 3.8 Flash and Flash-Lite TTS with voice cloning from 30-second samples, 2,000+ voices, and #1 spots on Hume AI benchmarks. Available today in the Gemini API.

Gemini 3.8 text-to-speech says hello
Gemini 3.8 text-to-speech says helloAI-generated
By Rebecca Stone5 min read

Updated

Why it matters

  • Gemini 3.8 Flash TTS scores #1 on Hume AI's Voice Design Benchmark (71.4) and leads accent modeling (60.8); Flash and Flash-Lite rank #1 and #2 on the Overall Quality Index.
  • Voice replication works from a 30-second audio sample, with consent verification, SynthID watermarking, and C2PA credentials built in.
  • Both models are available today in the Gemini API and Google AI Studio, in Gemini Notebook and Google Vids for consumers, and coming soon to Gemini Enterprise.

Google has launched two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, turning voice generation from a menu of static presets into what the company describes as a "dynamic creative studio" for creators, developers, and enterprises.

The release positions Google directly against rivals in a voice AI market where expressive quality, multilingual reach, and consent safeguards have become the deciding factors for buyers. It also extends a fast-growing Gemini Audio portfolio that now includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.

Two models, two jobs

Gemini 3.8 Flash TTS targets deep creative direction and character design. Users can build entirely new voices from scratch using natural-language prompts, then direct performances line by line with control over acting cues, pacing, dialect shifts, and backchanneling. Google pitches it at gaming, immersive audiobooks, podcasts, and interactive media.

Gemini 3.8 Flash-Lite TTS is the cost-efficient sibling, built for high-volume dubbing, audio content creation, and expressive voice agents, with fine-grained control over tone, pacing, and expressive nuance.

From 30 voices to an open-ended library

The headlining capability is generative voice design. With Gemini 3.8 Flash TTS, users can customize role, accent, and voice characteristics across more than 100 languages and dialects through natural-language prompting — Google's examples range from "a dramatic, fire-breathing dragon" to "a charismatic narrator with a distinct regional cadence."

The move scales Google's offering up from 30 original voices. A library of more than 2,000 production-ready voices is also available, with coverage of regional varieties including Mexican Spanish, Quebec French, and Scots English.

Voice replication is the most sensitive feature. The system can recreate a consistent vocal profile from just a 30 seconds of audio from the user's own voice or a voice they have the rights to use. Google says the feature ships with built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and vocal talent.

Users can save and manage custom voices to keep performance consistent with minimal drift across ongoing projects. A voice remixing capability is coming soon: users will be able to take a voice from the library and fine-tune timbre, pitch, pace, and accent with prompts such as "add subtle Southern US accent" or "soften the delivery."

Line-by-line direction

Both models give creators script-level control. Users can write their own stage directions or let Gemini steer delivery from natural script cues, spanning use cases from "a calm customer service agent to a whispered suspense scene."

Long-form generation holds voice quality, pacing, and character timbre across hours of continuous audio with minimal speaker drift — a core requirement for podcasts and audiobooks. Native two-speaker scene staging lets users direct multi-turn conversations from a single script while keeping both voices distinctly separated, with natural conversational turn-taking.

The models also support scripted vocal bursts and backchanneling. Non-verbal cues like <laughs>, <sigh>, and <gasp>, plus active-listening interjections like |mhm| or |yeah|, add conversational texture and what Google calls "precise comedic timing and reaction beats."

Benchmark claims

Google says Gemini 3.8 Flash TTS took the #1 overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and also leads in accent modeling at 60.8. On Hume AI's Overall Quality Index, the Flash and Flash-Lite models rank #1 and #2 respectively.

The company reports major improvements over Gemini 3.1 Flash TTS across long-form content and dual-speaker screenplay control. In blind human preference evaluations on Voice Arena, both models secure top positions among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

Consent and watermarking

Voice cloning has drawn regulatory and industry scrutiny, and Google built guardrails into the release. Before a replicated voice can be created, users must provide a verbal consent recording from the voice owner that matches the reference speaker.

Every audio clip generated by the Gemini Audio models carries a SynthID watermark, woven imperceptibly into the audio output. Google says the watermark keeps AI-generated speech detectable and helps prevent misinformation. The company has published a model card detailing its safety and responsibility approach.

Availability and partners

Developers can try the models today in Google AI Studio's audio playground, which works as a voice design workspace: prompt a new vocal identity from scratch or replicate your own voice, then move into a dual-speaker screenplay editor to direct line-by-line delivery.

Through the Gemini API, developer platforms including Agora, LiveKit, Pipecat, and Vercel support building and deploying speech generation experiences. Google is also partnering with Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, which are integrating the new TTS models to accelerate global dubbing, localize media with nuanced regional accents, and power conversational voice agents at scale.

Rollout starts today for Gemini 3.8 Flash TTS in the Gemini API and Google AI Studio for developers, with enterprise availability "coming soon" via API in Gemini Enterprise, and consumer availability in Gemini Notebook. Gemini 3.8 Flash-Lite TTS follows the same developer and enterprise path, and ships for everyone in Google Vids.

With benchmark leadership claimed, consent-verified cloning, and a 2,000-voice library in place at launch, Google has set the terms competitors in expressive TTS will now be measured against — and the pending voice remixing feature signals the company intends to keep expanding the creative controls.

Original: notebook.google.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

158 articles

Related articles

  1. Google's Flash TTS models build AI voices from text descriptions
  2. Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
  3. Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
  4. Google Ships Gemini 3.8 Live and Extended Thinking Models
  5. Google Launches Gemini 3.5 Live Translate Across Products

« Previous articleNext article »