Gemini 3.8 Live with Live Avatar gives Google’s AI a face
AI assistants have spent years getting better at understanding what we say. Google’s latest Gemini update changes the visual side of that interaction: instead of a voice coming from an invisible chatbot, Gemini 3.8 Live can now appear as a responsive digital character that talks, listens, sees, and keeps the conversation moving.
Google announced Gemini 3.8 Live with Live Avatar on September 24, 2026. The company describes it as a near-real-time visual presence for conversational AI. The important distinction is that this is not simply an animated talking-head video added after an answer is generated. Live Avatar is connected to Gemini 3.8 Live’s real-time dialogue system, allowing speech and video to be produced together during an active interaction.
Key Takeaways
- Gemini 3.8 Live with Live Avatar is currently available through Gemini Enterprise and the Gemini Live API.
- The avatar can generate synchronized 24 FPS video with speech, including lip movement and facial expressions.
- Gemini can continue talking while API, CRM, or ERP calls run in the background through asynchronous tool calling.
- Google says Live Avatar can switch between 97 languages while adapting speech, lip-sync, and expressions.
- Custom avatars are more restricted: Google says custom avatar creation is available through enterprise allowlisting and verification.
What is Gemini 3.8 Live with Live Avatar?
Gemini 3.8 Live is Google’s low-latency, real-time conversational model, while Live Avatar adds a generated visual persona to that interaction. In practical terms, developers can build an agent that responds with synchronized voice and video instead of voice or text alone.
Google’s documentation describes gemini-3.8-live as a model designed for bidirectional streaming. It accepts text, images, audio, and video and supports function calling, search grounding, interleaved reasoning, and the Live API. Google’s enterprise documentation adds 24 FPS Live Avatar video output as a new capability compared with earlier Live API generations.
This makes the feature especially interesting for customer-facing AI. A hotel concierge, insurance intake agent, virtual tutor, or product guide can maintain a visible presence while performing actions behind the scenes.
Why the avatar matters more than it sounds
The biggest change is not that Gemini can display a face. The more significant shift is that the face is part of a live conversational system.
Traditional AI video workflows usually separate the language model, text-to-speech system, avatar animation, and business tools. Each layer can introduce latency. Google is trying to make those pieces behave more like one continuous conversation.
That matters when a user interrupts, changes direction, points a camera at something, or asks the agent to complete a task. The goal is less “AI video” and more “AI agent with a visible presence.”
Key features worth knowing
Gemini 3.8 Live with Live Avatar combines several capabilities that become more useful when they operate simultaneously.
1. Real-time lip-sync and expressions
Google says Live Avatar generates synchronized video at 24 frames per second. The avatar’s facial expressions and lip movements are synchronized with synthesized speech, making the interaction feel more like a live conversation than a pre-rendered clip.
2. Asynchronous tool calling
This is one of the more practical features. Gemini can call APIs and business systems in the background while the conversation continues. Google gives a hotel check-in example: the agent can continue talking while a backend system retrieves or updates information.
For businesses, this can reduce the awkward silence that often appears when a voice bot has to wait for a database or external API.
3. Live visual understanding
Gemini 3.8 Live can process live camera feeds and screen shares alongside audio. That opens use cases where the user needs to show the AI something rather than describe it manually.
Google demonstrates an insurance claims workflow in which a customer can talk to an agent while showing damage on camera. The conversation can feed information into the claims process while other agents and tools work in the background.
4. Multilingual conversations
Google says Live Avatar can automatically detect and transition between 97 supported languages. The company says lip-sync and expressions adapt during those language changes without introducing visual drift.
For global customer-service deployments, this could be more useful than simply translating text because the same visual agent can remain present while the spoken language changes.
5. Custom avatar support
Organizations can use Google’s preset avatar library or create custom avatars from reference material. Google says custom avatar creation is currently restricted to enterprise allowlisting and verification, so this should not be confused with an unrestricted consumer avatar generator.
How Gemini Live Avatar works in practice
For developers, the basic workflow is straightforward: choose Gemini 3.8 Live, enable Live Avatar, select an avatar and voice, then configure the agent’s instructions and tools.
- Select the model: In Google Cloud’s Agent Platform Studio, select
gemini-3.8-live. - Enable Live Avatar: Choose the Live Avatar option in the real-time streaming interface.
- Select a visual persona: Choose a supported preset avatar, or use a custom avatar if your organization has access.
- Choose a voice: Pair the avatar with a supported voice.
- Add instructions: Define the agent’s behavior, tone, task boundaries, and business rules.
- Connect tools: Add APIs or business systems so the agent can take action during the conversation.
Google’s developer documentation says the Live API uses bidirectional WebSockets for streaming interactions. The model can receive audio, video, and text while producing audio and, with Live Avatar enabled, synchronized video.
Pricing and availability
There are two different availability questions here: the Gemini 3.8 Live model and the Live Avatar enterprise feature.
Availability caveat
As of September 25, 2026, Google says Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise. Custom avatar creation is gated behind enterprise allowlisting. This is not the same as saying the consumer Gemini app has received the same avatar experience.
| Gemini 3.8 Live API item | Current published standard price |
|---|---|
| Text input | $0.75 per 1 million tokens |
| Audio input | $3.00 per 1 million tokens or $0.005/minute |
| Image/video input | $1.00 per 1 million tokens or $0.002/minute |
| Text output | $4.50 per 1 million tokens |
| Audio output | $12.00 per 1 million tokens or $0.018/minute |
Google also lists a free tier for Gemini 3.8 Live in its Gemini API pricing documentation. Enterprise deployments can have additional requirements around provisioned throughput, compliance, endpoints, and data governance. Actual production cost will depend heavily on session length, modality, context accumulation, and tool usage.
What businesses can actually build with it
The strongest use cases are situations where seeing and hearing the agent provides a real benefit rather than simply making a chatbot look more human.
Customer service
A company could deploy a visual support agent that explains troubleshooting steps, watches what a customer is showing through a camera, and calls internal systems without ending the conversation.
Hotels and travel
A virtual concierge could answer questions, check availability, call booking systems, and guide guests through services while maintaining a visible persona.
Insurance and claims
Live video understanding is particularly relevant to claims intake. A customer can show damage while explaining what happened, while the AI extracts information and passes it into downstream workflows.
Training and education
A responsive avatar could act as a tutor or instructor, especially for demonstrations where the learner needs to speak, show a screen, or receive immediate feedback.
Interactive kiosks
Google specifically describes support for web, mobile, and interactive kiosk experiences. That gives businesses a path toward physical customer-service interfaces that can see, hear, speak, and execute actions.
Where Gemini Live Avatar still falls short
The technology is impressive, but the current availability model limits who can use it. Live Avatar is primarily an enterprise/developer capability, so consumers should not assume that the same experience is already built into every Gemini chat.
There is also a practical cost consideration. Real-time multimodal sessions can become expensive as context accumulates, particularly when audio remains active for long periods. Google’s Live API guidance explains that persistent sessions can re-process accumulated context on subsequent turns, making context management important for production applications.
Custom avatar access is another limitation. Google says organizations need enterprise allowlisting and verification for custom avatar creation. That is a sensible control for identity and impersonation risks, but it also means the most brand-specific version of the technology is not an open feature for everyone.
Finally, a face does not automatically make an AI agent more useful. In many software workflows, text or voice remains faster and less distracting. The visual layer makes the most sense when facial presence, visual guidance, or physical interaction contributes something meaningful.
Gemini 3.8 Live vs. a traditional voice AI
The main difference is modality. A conventional voice agent can hear and speak, while Gemini 3.8 Live can combine audio with live visual input and tool execution. Live Avatar adds synchronized video output, turning the agent into a two-way visual interface.
| Capability | Traditional voice bot | Gemini 3.8 Live + Live Avatar |
|---|---|---|
| Speech conversation | Yes | Yes |
| Live visual input | Usually limited | Yes |
| Animated visual persona | Usually no | Yes |
| Asynchronous tool calls | Depends on platform | Yes |
| Multilingual switching | Depends on platform | Google says 97 languages |
| Custom avatar | Varies | Enterprise allowlisting |
Safety and the problem of making AI look human
A more expressive AI also creates a bigger trust problem. When an avatar looks and sounds natural, users may find it harder to distinguish synthetic media from a real person.
Google says generated audio and video from its AI products carry imperceptible SynthID watermarks. The company positions this as a way to help identify AI-generated content and reduce misinformation and misattribution.
For businesses, that means responsible deployment should include clear disclosure that the user is interacting with an AI agent, appropriate identity controls, and restrictions around sensitive or impersonation-prone applications.
Who should actually pay attention to Gemini 3.8 Live Avatar?
Developers building real-time agents should pay attention because Google is combining several pieces that previously had to be assembled separately: speech-to-speech dialogue, visual input, tool execution, multilingual conversation, and generated video presence.
Businesses should look at it when the interaction itself matters. Customer service, hospitality, insurance, education, retail kiosks, and guided experiences are more obvious fits than ordinary internal chat.
For everyday consumers, the practical question is different: when will this experience become broadly available in the Gemini app? Google’s September 24 announcement specifically points to Gemini Enterprise availability, so the current release should be treated as an enterprise/developer launch rather than a general consumer Gemini feature.
FAQ
What is Gemini 3.8 Live with Live Avatar?
Gemini 3.8 Live with Live Avatar is Google’s real-time conversational AI system that combines speech, visual understanding, tool execution, and a synchronized animated video persona. Google says the avatar can generate 24 FPS video with speech-synchronized lip movements and expressions for interactive enterprise experiences.
Is Gemini 3.8 Live Avatar available to everyone?
No. As of September 25, 2026, Google says Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise and through its developer platform. Custom avatar creation is restricted to enterprise allowlisting, so the feature should not be treated as a standard consumer Gemini app capability.
How many languages does Gemini Live Avatar support?
Google says Live Avatar can automatically detect and transition across 97 supported languages. The system is designed to adapt speech, lip-sync, and facial expressions during language changes, allowing a visual agent to maintain its presence while the spoken language changes.
Can Gemini Live Avatar use tools while talking?
Yes. Gemini 3.8 Live supports asynchronous function calling, allowing an agent to execute API, CRM, ERP, or other tool calls in the background while the conversation continues. This is designed to reduce dead air while the AI waits for backend operations to finish.
Can I create a custom Gemini Live Avatar?
Google supports custom avatar creation from reference material, but access is restricted. Google says custom avatar creation is currently available through enterprise allowlisting and verification. Developers can also use preset avatars from Google’s supported library when custom avatar access is not available.
Final take
Gemini 3.8 Live with Live Avatar is less about giving a chatbot a pretty face and more about changing the interface for real-time AI agents. The combination of synchronized video, speech, live visual understanding, background tool execution, and multilingual interaction gives developers a much broader interface to work with.
The most interesting part is what happens behind the avatar. If an AI can keep talking naturally while checking a database, understanding a camera feed, switching languages, and completing a business transaction, the avatar becomes a visible front end for an agent rather than a novelty.
For now, the enterprise focus matters. Consumers should not expect the full Live Avatar experience inside the regular Gemini app simply because the Gemini 3.8 Live model exists. But for companies building customer-facing AI, Google has made a significant move toward AI agents that do not just answer you—they appear to be there with you.
An AI researcher who spends time testing new tools, models, and emerging trends to see what actually works.