Models

Google Rolls Out Gemini 3.8 Live with Live Avatar for Enterprise

Gemini 3.8 Live with Live Avatar ships in Gemini Enterprise, pairing real-time video personas with 97-language lip-sync, background tool calling, and SynthID watermarking.

Introducing Gemini 3.8 Live with Live Avatar
Introducing Gemini 3.8 Live with Live AvatarAI-generated
By Elena Vasquez4 min read

Updated

Why it matters

  • Gemini 3.8 Live with Live Avatar launched in Gemini Enterprise, one week after the Gemini 3.8 Live launch.
  • Live Avatar supports seamless transitions across 97 languages with adaptive lip-sync and no visual drift, per Google.
  • All AI-generated audio and video output is watermarked with SynthID; custom avatar creation requires enterprise allowlisting.

Google has launched Gemini 3.8 Live with Live Avatar, a feature that pairs near real-time video generation with speech to give the company's live dialogue models a dynamic visual persona that "listens, sees, and speaks." The feature is available starting today in Gemini Enterprise, one week after the launch of Gemini 3.8 Live.

The release puts a concrete product behind one of the most contested fronts in enterprise AI: real-time multimodal interfaces. Companies deploying customer-facing agents have so far had to choose between voice-only assistants and pre-rendered video characters that break the illusion of conversation. Google is betting that precise lip-syncing, natural expressions, and fluid turn-taking — the three capabilities the company highlights — will let enterprises expand virtual offerings such as customer service and interactive walkthroughs into what the announcement calls "richer, more accessible experiences."

Multimodal by design

Google frames the feature around a simple observation about human conversation. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," the company states. Live Avatar brings those capabilities to enterprise agents by processing visual and audio inputs simultaneously and responding with expressive audio and video in near real time.

The company demonstrates the system taking in what it sees and hears and replying with synchronized speech and facial animation. According to Google, this simultaneous processing generates "enriching conversations for a more comprehensive experience." The claim matters because real-time visual understanding combined with real-time video generation is computationally demanding; most avatar products on the market today animate a static character over generated speech rather than generating the video stream itself.

Reasoning underneath the face

Live Avatar is more than a talking head, according to Google. The feature is backed by Gemini's advanced reasoning and supports asynchronous tool calling. In practice, the avatar can trigger tool calls and fetch data in the background while the conversation continues — the announcement describes the system "handling complex tasks while ensuring an uninterrupted conversational flow."

Google's example is a hotel guest check-in, where the avatar calls tools in the background while the dialogue continues without pause. For enterprises, this addresses a familiar failure mode of voice agents: the dead air that follows when an assistant pauses to query a reservation system or database. Whether Google's implementation holds up under real-world latency remains to be tested, but the architecture — decoupling tool execution from the conversational turn — is the same design pattern driving agentic AI deployments across the industry.

97 languages, one face

The most technically aggressive claim in the announcement concerns multilingual support. Live Avatar features what Google calls native multilingual speech-to-speech synchronization, dynamically adapting lip-sync and expressions as it "can seamlessly transition across 97 languages without degrading video fidelity or introducing visual drift."

Google shows the avatar switching between languages mid-conversation, with lip-sync and expressions adapting on the fly. Cross-language lip-sync is a hard rendering problem: mouth shapes differ across languages, and naive systems produce visible artifacts when the spoken language changes. If the fidelity claim holds, it would remove a significant barrier for global enterprises running a single conversational agent across markets.

Custom avatars, gated by allowlist

Brand identity is the other enterprise requirement Google addresses. The product ships with a library of diverse preset avatars, each with what the company describes as "a distinct look, voice, and expressive presence." Organizations can also build their own. From a high-quality reference image, developers can generate a fully animated, responsive avatar that preserves reference likeness, brand styling, or character identity.

There is a governance catch: custom avatar creation is currently available only through enterprise allowlisting. The restriction signals that Google is controlling access to likeness generation — the functionality most open to abuse — while letting the preset library circulate freely.

Watermarks and identity safeguards

Google says it built Live Avatar with "strict safeguards designed to respect identity, and keep AI-generated content transparent." All output generated by Google's AI products carries a SynthID watermark, an imperceptible marker woven directly into the audio and video output. According to the company, the watermark helps ensure AI-generated content remains detectable, minimizing misinformation and misattribution. Google directs enterprises to its model card for its full approach to safety and responsible deployment.

The transparency layer is not incidental. Photorealistic, real-time talking avatars are a known vector for fraud and impersonation, and regulators including the EU under the AI Act are pushing for provenance marks on synthetic media. Embedding detection at the output level — rather than relying on platform labeling — is the approach Google has standardized on across its generative products.

Availability and what to watch

Gemini 3.8 Live with Live Avatar is available now in Gemini Enterprise, with API documentation published for developers getting started.

The launch sequence itself is notable: Gemini 3.8 Live shipped one week ago, and the avatar layer follows immediately. Google is compressing the gap between its base dialogue models and the interfaces enterprises actually deploy. The next questions are pricing at scale — the announcement does not list any — and whether the 97-language lip-sync and uninterrupted tool calling survive contact with production workloads, which will determine whether Live Avatar becomes a default for enterprise customer-facing AI or a premium demo.

Original: blog.google

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. Google Gives Gemini 3.8 Live Agents Talking Avatars
  2. Google's Live Avatar Puts a Human Face on Gemini Enterprise Agents
  3. Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
  4. Google Launches Gemini 3.5 Live Translate Across Products
  5. Google Ships Gemini 3.8 Live and Extended Thinking Models

« Previous articleNext article »