Models

Google Launches Gemini 3.5 Live Translate Across Products

Gemini 3.5 Live Translate detects 70+ languages, streams speech continuously, and ships today into Meet, the Google Translate app, and the Gemini Live API for developers.

Fluid, natural voice translation with Gemini 3.5 Live Translate
Fluid, natural voice translation with Gemini 3.5 Live TranslateAI-generated
By Elena Vasquez4 min read

Updated

Why it matters

  • Gemini 3.5 Live Translate detects 70+ languages and translates speech continuously, staying just a few seconds behind the speaker while preserving intonation, pacing, and pitch.
  • Google Meet speech translation expands from five languages and English-only routing to 70+ languages and over 2,000 language combinations per meeting, in private preview this month.
  • Grab is testing the model on voice calls between drivers and travelers, a flow that handles over 10 million voice calls per month; all generated audio carries a SynthID watermark.

Google has released Gemini 3.5 Live Translate, an audio model for live speech-to-speech translation that automatically detects 70+ languages and preserves a speaker's intonation, pacing, and pitch. The model is available starting today: in public preview for developers through the Gemini Live API and Google AI Studio, in private preview this month for enterprises in Google Meet, and globally for consumers in the Google Translate app on Android and iOS.

The release marks a structural break from conventional translation pipelines. Turn-by-turn systems wait for the speaker to finish before responding. Gemini 3.5 Live Translate generates speech continuously, balancing what Google describes as "the trade-off between waiting for context to improve quality and translating immediately to stay in sync with the speaker." The result is audio without awkward pauses, trailing the speaker by just a few seconds for the duration of a session.

The stakes are significant. Google says it now translates over a trillion words each month for billions of users across its products, a line of business that started twenty years ago as one of the company's early machine learning experiments. Real-time speech translation is the next competitive frontier, and Google is pushing it into developer tools, enterprise meetings, and a consumer app in a single coordinated launch.

For developers: streaming, multilingual, noise-robust

The model processes speech as it streams and handles multilingual inputs without manual configuration. Google also emphasizes noise robustness, meaning applications built on the model can operate in loud, unpredictable environments. The company points to use cases including live interpretation for multilingual calls, meetings, lessons, and broadcasts.

Developers can access a demo and example code in the Gemini Cookbook. Real-time media streaming is handled through the Gemini Live API integrations with developer platforms including Agora, Fishjam, LiveKit, Pipecat, and Vision Agents. Google says these integrations manage the streaming infrastructure so developers can focus on user experience.

One partner is already testing the model at scale. Grab, the Southeast Asian platform, is using Gemini 3.5 Live Translate to enable near real-time multilingual communication between drivers and travelers at pickups. Those users make over 10 million voice calls per month through Grab, according to Google. Companies including CJ ENM and LiveKit have also shared positive early feedback, highlighting the model's translation quality, accuracy, and low latency.

Google Meet jumps from five languages to 70+

The enterprise rollout carries the most concrete before-and-after numbers. Speech translation in Google Meet, powered by 3.5 Live Translate, will support 70+ languages, up from a previous limit of just five. Meetings will support over 2,000 language combinations in a single session, expanding from a prior state in which Meet only translated to and from English. The interface is also being updated to give instant access to speech translation.

The update launches in private preview for select business Google Workspace customers starting this month, with a broader rollout planned later this year.

Consumer features and a new listening mode

In the Google Translate app, the Live translate feature now works with any pair of connected headphones, delivering translation that mirrors the speaker's tone across 70+ languages.

Android users get an additional capability: a new "listening mode" that streams translated audio directly through the phone's earpiece. Users hold the phone to their ear as they would on a regular call. Google positions the feature for situations where someone wants to hear a translation privately without headphones, and gives a concrete example: hearing a near real-time English translation of a guided tour in Spanish directly through the earpiece.

SynthID watermarking on all output

All audio generated by the model carries a SynthID watermark, woven imperceptibly into the audio output. Google says the watermark keeps AI-generated content detectable to help prevent misinformation. The company has published a model card detailing its safety and responsibility approach.

With today's launch, Google has collapsed the gap between experimental speech translation and shipped product across three market segments at once. The broader Workspace rollout later this year, and Grab's testing at 10 million monthly voice calls, will show whether continuous translation holds up at production scale.

Original: blog.google

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
  2. Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
  3. Google Ships Gemini 3.8 Live and Extended Thinking Models
  4. Google Ships Gemini 3.1 Flash Live Audio Model Globally
  5. Google Rolls Out Gemini 3.8 Live with Live Avatar for Enterprise

« Previous articleNext article »