Models

Google Ships Gemini 3.8 Live and Extended Thinking Models

Google's Gemini 3.8 Live Extended Thinking tops Artificial Analysis' Speech to Speech Quality Index at 82.6, as both new models roll out to developers, enterprises, and consumers starting today.

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Introducing Gemini 3.8 Live and 3.8 Live Extended ThinkingAI-generated
By Elena Vasquez3 min read

Updated

Why it matters

  • Gemini 3.8 Live Extended Thinking ranks #1 on Artificial Analysis' Speech to Speech Quality Index with 82.6, and scores 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking.
  • Gemini 3.8 Live supports 97 languages with automatic mid-conversation transitions and executes tool calls in the background while continuing to talk.
  • Both models are available starting September 17, 2026 in the Gemini API, Google AI Studio, private previews of Gemini Enterprise, Search Live, and consumer Gemini apps.

Google has released two new speech-to-speech models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed to push near real-time reasoning into production voice agents and everyday conversational AI.

The announcement, updated September 17, 2026, positions the pair at different ends of the market. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking targets high-complexity tasks with increased intelligence and multi-step reasoning.

The stakes are considerable. Voice is becoming the primary interface for enterprise agents in customer service, banking, and workflow automation, and Google is competing with other frontier labs on both benchmark performance and price for developers building production systems.

Benchmark results

Google says Gemini 3.8 Live Extended Thinking took the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It also leads in agentic task completion, scoring 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark, while posting 97.7% on Big Bench Audio for reasoning. The company says the model maintains a highly competitive price point compared to other frontier models.

Gemini 3.8 Live secured second place in the Speech Agent Arena and remains highly cost-effective, according to Google, making it a capable model built for scale.

On ServiceNow's EVA-Bench, a benchmark for evaluating voice agents, Google says its models push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality. The company notes these runs were performed on the Live API on the Gemini Enterprise Agent Platform.

Real-time behavior

Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with additional context. It automatically detects and transitions between 97 supported languages mid-conversation. The model executes tools and API calls in the background while continuing to talk, so it can acknowledge a request and keep chatting while tasks finish.

Extended Thinking reasons and speaks simultaneously. For multi-step background tasks, it uses early verbal cues like "Let me check that…" to acknowledge prompts naturally, and provides live progress narration to walk users through tasks as they progress — all while maintaining an uninterrupted conversational flow.

Google says the Live models deliver more intuitive, collaborative experiences across Google Workspace, Search, and the Gemini app, especially for complex tasks handled entirely by voice.

Ecosystem and partners

Through the Gemini Live API, developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents let developers build and deploy high-performance voice-driven interfaces. These platforms manage the real-time media streaming infrastructure, freeing developers to focus on user experience.

Google is also partnering with Salesforce, Genspark, and Lumeris, which the company says highlighted the models' latency, fluidity, and tool-calling capabilities.

All audio generated by Google's AI products carries a SynthID watermark woven imperceptibly into the output, keeping AI-generated content detectable to help prevent misinformation. Google directs readers to the model card for details on its safety approach.

Availability

Gemini 3.8 Live is rolling out starting today in the Gemini API and Google AI Studio for developers, in private preview in Gemini Enterprise for enterprises — with Gemini Enterprise for Customer Experience coming soon — and in Search Live for all users.

Gemini 3.8 Live Extended Thinking follows the same developer and enterprise rollout, and is additionally coming soon to Google Workspace business customers. Consumers get it in Gemini Live, and for Google AI Pro and Ultra subscribers in Workspace in Docs, plus all Google AI subscribers in Gmail and Keep.

With a top-ranked speech-to-speech model, multi-step reasoning that runs while the user is still talking, and background tool execution across 97 languages, Google is betting that voice agents — not chat windows — become the default way enterprises and consumers delegate complex work.

Original: artificialanalysis.ai

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

147 articles

Related articles

  1. Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
  2. Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
  3. Google Ships Gemini 3.1 Flash Live Audio Model Globally
  4. Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
  5. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month

« Previous articleNext article »