Google is extending its Gemini 3.8 Live model with "Live Avatar": animated avatars that hold near-real-time conversations with users. Speech, lip movement and facial expression are generated in sync, and the avatars can process what users show them through a camera or screen share and react to it in the conversation.
The system supports 97 languages and can switch between them automatically mid-conversation, adapting lip movement and facial expression to each language. It is designed to keep the conversational context even when users interrupt. In parallel, the agent can call tools and interfaces — querying data or completing bookings — without breaking the flow. Google demonstrated this with a hotel check-in in which the avatar kept talking while information was fetched in the background.
Live Avatar launches exclusively as an enterprise feature and is generally available now through Gemini Enterprise. Customers can choose from a library of prebuilt avatars with different appearances, voices and manners. In one demonstration Google created a custom avatar from a reference photo and an audio recording, preserving the likeness, brand style or character identity of the source — a capability that requires explicit activation and verification by Google to prevent identity abuse. All generated audio and video streams carry Google's imperceptible SynthID watermark.
Businesses can deploy the avatars on websites, mobile devices and interactive kiosks; Google names digital concierges and customer service as likely uses. No consumer version has been announced. The basis is Gemini 3.8 Live, presented the previous week as a multimodal model for real-time voice dialogue, handling audio, image and video input. In Artificial Analysis's Speech-to-Speech Quality Index, the more capable extended-thinking variant takes first place with 82.6 points among all tested real-time voice models; Gemini 3.8 Live ranks fifth.



