Products & Tools

Google's Gemini 3.5 Transcribe posts 2.6% word error rate

Google's Gemini 3.5 Transcribe debuts with a 2.6% WER in non-streaming and 4.0% in streaming on Artificial Analysis tests, a 70% latency cut versus Chirp 3, and support for 85 languages across two API endpoints.

By Rebecca Stone4 min read

Updated

Why it matters

  • Gemini 3.5 Transcribe posts a 2.6% non-streaming and 4.0% streaming WER in Artificial Analysis testing.
  • Time to final transcription improves 70% versus Google's prior Chirp 3 model.
  • The FLEURS multilingual benchmark scores 5.50% WER (streaming) and 5.04% WER (non-streaming).
  • The model auto-detects and transcribes more than 85 languages via the Live and Interactions APIs.
  • Six voice-agent platforms—Agora, Fishjam, LangChain, LiveKit, Pipecat, and Vercel—plus Vision Agents have integrated the Live API on day one.

Google's Gemini 3.5 Transcribe posted a 2.6% word error rate (WER) on non-streaming audio and 4.0% on streaming audio in independent testing by Artificial Analysis, the company said today as it opened public preview access to the new speech-to-text model through the Gemini API.

The model also cuts time-to-final-transcription by 70% compared with Google's prior transcription engine, Chirp 3. That latency gain matters for real-time voice agents, where a half-second slowdown can break conversational turn-taking.

"Today, we're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions," Google wrote in the launch announcement.

What does 3.5 Transcribe do differently from a conventional transcriber?

Google pitched the model as a step beyond raw speech recognition. Where a vanilla ASR engine returns a literal transcript, 3.5 Transcribe cleans up the speaker on the fly:

  • Self-corrections resolve mid-stream ("let's meet Tuesday—no, Wednesday").
  • Filler words ("ums," "ahs") drop out.
  • Punctuation and formatting appear without post-processing.
  • Numeric strings, postal codes, and order IDs survive noisy conditions intact.

The result lands closer to a polished draft than a raw transcript, a useful shift for meeting notes, customer-service logs, and any pipeline that feeds downstream language models.

What do the benchmarks show?

Third-party numbers from Artificial Analysis place 3.5 Transcribe in the top tier of public speech recognizers:

  • Average WER of 4.0% in streaming mode and 2.6% in non-streaming mode.
  • 70% improvement in time to final transcription versus Chirp 3.
  • FLEURS multilingual benchmark: 5.50% WER streaming, 5.04% WER non-streaming across a top-languages set.

The FLEURS result matters for non-English markets. Chirp 3 lagged on multilingual tasks; the new model closes that gap, per Google's numbers, and ships with automatic language detection across more than 85 locales, including regional accents and dialects.

How are developers integrating it?

Google shipped 3.5 Transcribe across two API surfaces, each tuned for a different workflow:

  • gemini-3.5-transcribe-live via the Live API: bidirectional streaming with sub-second latency for interactive voice apps.
  • gemini-3.5-transcribe via the Interactions API: post-process transcription with speaker attribution and word-level timestamps for recorded audio, meetings, and call logs.

Voice-agent platforms moved quickly to add support. Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents all expose the model to their customers through the Live API, Google said. LiveKit and Pipecat are widely used open-source frameworks for production voice agents, and LangChain's LangGraph handles tool routing; shipping 3.5 Transcribe inside those stacks shortens the path from prototype to a deployed customer-service bot.

Which Google surfaces use it today?

The model is no longer a developer-only preview. It already powers consumer features rolling out now:

  • Rambler on Android, a Gboard dictation mode that strips filler words and accepts voice commands to retone or correct text.
  • The Gemini app on macOS, which transcribes free-form speech into formatted text and routes voice commands to other Gemini models for image generation and file summarization.
  • Google Antigravity, where the model pairs screen and chat context, with user permission, to keep file names and agent states spelled correctly inside active documents.
  • Google AI Studio's Build mode, where voice-driven app assembly has been added for hands-free prototyping.

Chrome dictation in any web text field is on the roadmap, Google said, without committing to a specific launch date.

What about function calling and customization?

Two capabilities distinguish 3.5 Transcribe from a plain transcription endpoint:

  • Function calling lets the model delegate side tasks such as image generation and file analysis to other Gemini models in the background. The feature is live in the Gemini macOS app today.
  • Custom vocabulary ingests a list of jargon, proper nouns, or product SKUs and steers the recognizer toward them, so drug names, brand handles, and technical terms stop getting mangled.

The Interactions API labels up to three speakers in pre-recorded audio with timestamps. Support for four or more speakers is labeled experimental.

Who is using it in production?

Early customers cite the model's latency, accuracy, and language support as decisive. They include smartphone maker vivo, healthcare transcription firm Intellitek Health, and Lingopal, a localization platform.

Why does the launch matter?

Speech-to-text has become the front door of conversational AI. Every voice agent, meeting note-taker, and live caption system routes through it. Cutting streaming WER below 4% and trimming latency by 70% against Chirp 3 makes common failure modes (mangled proper nouns, awkward pause-then-burst pacing) rarer, though jargon and overlapping speakers still trip every ASR engine at this maturity level.

The integration story is what Google is selling, not just the benchmark chart. With Chirp 3, developers had to bolt on cleanup models for fillers and punctuation. With 3.5 Transcribe, that polish ships inside the same endpoint, and function calling carries the pattern into tool-driven workflows on consumer devices. As LiveKit, LangChain, and the rest of the agent-framework ecosystem plug into the model, the competition will tilt from raw accuracy toward how tightly each provider couples transcription to downstream tool calls, the layer Google is now banking on.

Original: aistudio.google.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

214 articles

Related articles

  1. Google Launches Gemini 3.5 Live Translate Across Products
  2. Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
  3. Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
  4. Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
  5. Google Ships Gemini 3.1 Flash Live Audio Model Globally

« Previous articleNext article »