Google Gemini Gets Live AI Avatars That Speak 97 Languages With Real-Time Lip Sync

Google has brought AI capabilities to the next level with its new Live Avatar feature for Gemini 3.8 Live. The feature lets users interact with an animated AI persona during a live conversation where speech, facial expressions and synchronized lip movements come together to create a more natural video experience. Google said its AI system automatically recognizes the language spoken and enables conversations in 97 languages.

Google Gemini | Photo Credit: https://gemini.google.com/
Google Gemini | Photo Credit: https://gemini.google.com/

The new functionality will help AI interactions be more visual and conversational and not only text or audio-based but also a digital persona, which responds to the user through video. Google says the avatar can also be programmed to change its lip movements and facial expressions with the language used, so that it can stay synchronized during multilingual conversations.

The Live Avatar feature is currently available on Gemini Enterprise and is available in US and European platforms. It follows Google's launch of Gemini 3.8 Live and Gemini 3.8 Live Extended, which are speech-to-speech AI models that can be used in low-latency conversations and agentic tasks. While Gemini 3.8 Live Extended is currently in private preview, Live Avatar is available for supported Gemini 3.8 Live environments.

Gemini Live Avatar Supports 97 Languages

One key feature of Google's new system is its multilingual capabilities. The Live Avatar can automatically detect the user's language and respond using speech while keeping the facial movements synchronized. Users can communicate with the AI in their own language without having to manually change the settings during the conversation.

Google says the system has been designed to dynamically adapt its lip synchronization and expressions as the conversation moves between languages. This is especially relevant for businesses that are looking to create customer-facing AI experiences for audiences speaking different languages.

The company is also providing a set of pre-built avatars for companies. Organizations can choose from the available digital personas or create their own customized avatars based on reference images, giving businesses a lot more freedom in creating branded AI assistants.

AI Can See, Speak And Process Information At The Same Time

The Live Avatar system is not limited to simply answering spoken questions. It can process audio and visual inputs simultaneously while performing other tasks in the background, Google said.

The AI can use a phone or laptop camera to continuously interpret visual information. It can also work with screen sharing, allowing the system to react to changes in the user’s screen without the user having to manually take and upload an image.

At the same time, the AI can make tool calls and access information in the background. Google says these processes can take place without interrupting the ongoing conversation, so interaction is continuous and responsive.

The technology can also be useful in situations where AI needs information about what a person is saying or what they’re showing. A product may be guided through by a system, information is shown on screen or a video-based interface.

Businesses Could Use Live Avatars For Customer Support

Google is positioning the technology as a potential platform for organizations developing interactive video agents. Companies could use customized AI personas for applications such as virtual customer-support representatives, AI concierges, information kiosks and other customer-facing services.

The ability to integrate real-time speech, visual understanding and facial expressions can lead to a more human interface than text-based chatbots. Businesses can also customize the look of their digital personas, potentially allowing them to create avatars that suit the brands or services of the company.

However, the technology is currently only for enterprise usage, so its early availability is aimed at organizations rather than general consumer users.

SynthID Watermarking Adds An AI Safety Layer

Google is also emphasizing content authenticity and safety as part of the Live Avatar rollout. Generated video content includes SynthID, Google's invisible digital watermarking technology developed by Google DeepMind.

The watermark is intended to help identify AI-generated material and to make synthetic content easier to distinguish from authentic recordings. This could become more significant as AI-generated video becomes more realistic and is used in customer-facing applications.

Google says the watermarking approach is intended to minimize misinformation and misattribution. It has also provided users with its model documentation for information about safety and responsible deployment.

With Gemini 3.8 Live’s Live Avatar feature, Google is moving beyond voice-based AI conversations to interactive multilingual video communication. Support in 97 languages and real-time lip synchronization, visual comprehension, custom avatars and SynthID watermarking could make the technology relevant to businesses experimenting with AI-based customer experiences.