📑 Table of contents

Gemini 3.8 Live with Live Avatar: Google adds real-time visual presence to its conversational assistant

LLM & Modèles 🟢 Beginner ⏱️ 15 min read 📅 2026-09-27

Gemini 3.8 Live with Live Avatar: Google adds real-time visual presence to its conversational assistant

🔎 Google's assistant has a face — and that face is generated by the model itself

On September 24, 2026, Google announced Gemini 3.8 Live with Live Avatar: its real-time conversational assistant now appears as a dynamic visual persona that listens, sees, and speaks. The official blog post describes a near-real-time presence, with precise lip-sync, natural expressions, and smooth turn-taking.

The pace is telling. Gemini 3.8 Live, the voice layer, launched on September 15, 2026. Nine days later, Google grafts video onto it. That same week, the company released Gemini 3.8 Flash TTS and Flash-Lite TTS, as reported by AndroidHeadlines. Voice, face, synthesis: Google is stacking the building blocks of an embodied assistant, one layer at a time.

My reading after reviewing the sources: the novelty isn't the avatar concept — synthetic faces have been around for years. It's the architecture. The video stream is produced directly by the model, not by an animation engine bolted on afterward. And access, for its part, is deliberately narrow: general availability in Gemini Enterprise (United States and European Union), nothing in the consumer app, nothing public on the standard API.


The essentials

  • Enterprise GA on September 24, 2026: Live Avatar is available in Gemini Enterprise on US and EU endpoints — not in the consumer Gemini app.
  • Video comes from the model: 24 FPS MP4 streams (response_modalities=["VIDEO"]), the face is produced by the model itself, with no separate avatar to animate.
  • 97 languages with native voice-face sync, mid-conversation language switching, with no degradation in video fidelity or "visual drift".
  • Asynchronous tool calling: the agent queries your systems while the conversation continues — official demos: hotel check-in and insurance claims processing.
  • Custom avatars locked: 1 reference photo + 1 audio sample are enough, but the option is reserved for enterprise customers allowlisted via Google Cloud sales; everyone else goes through a library of pre-designed avatars.
  • Invisible SynthID applied to all generated audio and video streams.
  • No model ID or public pricing on the standard API; Gemini 3.8 Live Extended Thinking remains in private preview.

Tool Main Use Price (September 2026) Ideal For
Gemini Enterprise + Live Avatar Live dialogue with a model-generated visual persona Quote-based (September 2026, check cloud.google.com) Customer support, AI video conferencing, kiosks
Gemini API — gemini-3.8-live Real-time voice without STT pipeline, asynchronous tools ≈ $1.38/hour (September 2026, check ai.google.dev) Developers: the voice component, available now
Agent Development Kit (ADK) Agents, session memory, audio streaming Open source Building the business logic behind the avatar
Gemini 3.8 Flash TTS / Flash-Lite TTS Cost-effective speech synthesis Pricing in the API docs (September 2026) Voice prototypes, simple agents

What Live Avatar changes compared to "classic" avatars

The break is architectural, not cosmetic. Until now, a conversational avatar was an assembly: a language model, a speech synthesis engine, and an animated face layered on top — a pre-modeled 3D head or an audio-driven reference video. Here, Google claims that the face is an output of the model itself, on par with the sound.

The three flaws of the two-stage pipeline

The classic architecture stacks up three problems. Latency, first: each stage (text → audio → animation) adds its own delay, and the conversation pays the price. Lip-sync, next: aligning an animation to a separately produced track yields approximate synchronization, especially in multilingual settings. Visual drift, finally: over long sessions, the animated face tends to degrade or drift away from its reference — the "visual drift" well known to avatar users.

The official announcement claims exactly the opposite: native multilingual voice-face synchronization across 97 languages, with no degradation in video fidelity and no visual drift. A marketing claim as long as no independent testing on long sessions has been published — but if it holds up, the category changes in nature.

The comparison in one table

Criterion Classic pipeline (TTS + animation) Live Avatar (24/09/2026)
Video origin Separate animation engine Generated by the model
Latency Accumulated at each stage Single loop, native speech-to-speech
Lip-sync Aligned to the audio track Native, multilingual (97 languages)
Visual drift Frequent on long sessions Announced as absent (to be verified)
Access Many providers GA enterprise only

My take: the demos — hotel check-in, insurance claim — are less interesting than this table. The real issue is the coupling of dialogue + action + video in a single generation loop.


Under the hood: a single loop for voice, vision, and video

Live Avatar is built on Gemini 3.8 Live, launched on September 15, 2026, and on years of work around Project Astra, as WebProNews recalled on September 25. Three technical choices deserve a closer look.

A native video stream, not an overlay. The developer requests the video modality (response_modalities=["VIDEO"]) and receives synchronized MP4 segments at 24 FPS, produced by the model. WindowsForum, which covered the enterprise GA in detail, insists on this point: the face is produced by the model itself, not by a separate avatar being animated.

Speech-to-speech without an STT pipeline. No intermediate transcription: audio goes in, the response comes out. The integration with the Agent Development Kit (ADK) makes it possible to define agents, session memory, and audio streaming without the traditional speech-to-text pipeline. Fewer layers means less latency — and fewer nuances lost along the way.

Asynchronous tool calls. This is the most important point for business use. While the avatar speaks, the agent queries your APIs. In the insurance demo, the agent checks policies and rules in the background while the video keeps playing. In the hotel demo, check-in moves forward while the customer chats. The conversation never pauses to wait for a fetch.

Add simultaneous vision — camera feed and screen sharing at the same time — and you get an interlocutor that sees what you're showing while responding to you. For the details of the audio layer, our analysis of Gemini 3.8 Live: Google's real-time voice at $1.38 per hour — background reasoning and tool execution without interrupting the conversation remains the reference.

Fabien Blanc-paques, group PM for Gemini Live, sums up the direction: "our priority has shifted to the quality of every interaction" (WebProNews, September 25, 2026). Translation: Google is opting for reliability, interaction by interaction, rather than a race for features. For a product that displays a generated human face, that's the right trade-off — an avatar that derails destroys more trust than an awkward text chatbot.


Availability and pricing: what is actually accessible today

Direct answer: Live Avatar is only accessible via Gemini Enterprise, on US and EU endpoints, with a contract. Neither the consumer app nor the standard API exposes the feature as of September 24, 2026.

Component Channel Status (September 2026)
Live Avatar (video + voice) Gemini Enterprise GA — US and EU, quote-based
gemini-3.8-live (audio only) Gemini API Available, ≈ $1.38/hr
gemini-3.8-live-extended-thinking Gemini API Private preview
Custom avatars Enterprise allowlist Via Google Cloud sales only
Flash TTS / Flash-Lite TTS Gemini API Released the week of September 24

The clarification from AINewsReport is essential: Live Avatar is not a separate model ID on the general Gemini API. The list of live models only shows gemini-3.8-live and gemini-3.8-live-extended-thinking, with no avatar pricing line. The target is the enterprise buyer and their integrators — not the individual developer.

It's frustrating, but consistent. A product that generates a conversation partner's face in real time raises questions of identity, consent, and fraud that Google clearly doesn't want to release as self-service. The allowlist via Google Cloud sales is a deliberate filter: the vendor chooses its first deployments.

For developers who want to work with the material now, the audio component remains accessible at about $1.38 per hour. Our roundup of free AI APIs also lists options for prototyping without blowing the budget.


Freelance and SMB Use Cases: Four Concrete Scenarios

Google's demos highlight four use cases, all transferable to SMBs: customer support with a synthetic face, AI video conferencing, interactive kiosks, and product walkthroughs. None of them fundamentally requires being a large enterprise — only product access is, for now, enterprise-level.

Customer Support with a Synthetic Face

This is the most polished scenario. In the insurance demo, the customer describes their claim on video; the avatar maintains a fluid conversation while the agent checks policies and rules in the background. Transposed to an SMB: a visual front end that absorbs call spikes, 24/7, connected to your business tools. The gain isn't that the face replaces a human — it's that the conversation doesn't get interrupted during verification checks.

AI Video Conferencing and Automated Reception

The hotel check-in demo shows an avatar greeting a customer and processing their arrival while they chat. Variations: reception kiosks in branch offices, interactive kiosks at airports or events, visual receptionists for small businesses without a secretarial team. The kiosk is probably the first accessible market: controlled environment, bounded scope, measurable ROI.

Product Walkthroughs and Training

An avatar that guides the user through software, sees the shared screen, and adapts to their language — 97 languages natively covered — changes the scale of onboarding. For a SaaS vendor, it's a multilingual product demo without a support team in ten countries. For a training provider, a visual tutor available around the clock.

The Real Opportunity for Freelancers: The Agent Layer

The most interesting point: AINewsReport points out that the target is the enterprise buyer and its integrators. The value won't be in the video — provided by Google — but in the agent behind it: business logic, tool calls, session memory. That's exactly what the ADK makes it possible to build today, and what we cover in our guide to the best LLMs for AI agents.

Another complementary read: AI Avatar + personal assistant: the ultimate productivity combo, for the individual version of the phenomenon — a face-based assistant that manages your calendar and follow-ups, not just your customer service.


The limitations to keep in mind

Three limitations, all documented or deducible from the sources, temper the enthusiasm.

Access, first. GA enterprise only, allowlist for custom avatars, Extended Thinking in private preview: almost everything interesting is behind a contract. Google publishes neither avatar model IDs nor public pricing on the standard API.

Lack of hindsight, next. The absence of "visual drift" across 97 languages is a claim from September 24, not a result of independent tests on hour-long sessions. Traditional pipelines degraded precisely on long sessions; that's the first criterion to check as soon as enterprise feedback starts circulating.

Cost and network, finally. No public pricing exists for video, but generating and streaming 24 FPS MP4 will mechanically cost more than audio — and will require a solid connection on the user's side. For a wired kiosk, no problem. For a poorly covered mobile client, test it before promising anything.

SynthID and guardrails: what Google enforces, and what remains to be done

All generated streams — audio and video — carry an invisible SynthID watermark, and access to custom avatars is locked behind an enterprise allowlist. That's the framework Google laid out at launch.

Two levels of guardrails emerge from the sources. The first is technical: SynthID, applied to every generated stream, makes it possible to identify a video or a voice as synthetic. The second is procedural: custom avatars — generated from a reference photo and an audio sample — are strictly reserved for allowlisted customers via Google Cloud sales, as WindowsForum points out. Everyone else goes through the library of pre-built avatars.

My take: it's necessary, but not sufficient. A watermark only protects if distribution platforms check for its presence. And the allowlist protects Google, not your users. If you deploy a synthetic face on the front line of customer support, explicitly display its artificial nature — in the interface, not at the bottom of a legal notices page. Trust is a product feature, not a legal cartel.

You're not an enterprise: how to prepare now

You won't be signing a Gemini Enterprise contract this week. But you can start building today the only layer that will matter when access widens: the agent. Three steps, in order of priority.

1. Make your agent speak. gemini-3.8-live is accessible via the API at roughly $1.38 per hour (September 2026, check ai.google.dev). It's the best learning ground: turn-taking, interruptions, async tools. Voice UX has its own laws — better to tame them on audio than on the day video arrives.

2. Orchestrate with the ADK. Agents, session memory, audio streaming: Google's open source kit covers exactly what Live Avatar expects beneath it. A well-designed agent today will plug into the visual layer tomorrow without a rebuild.

3. Choose the brain and prepare the front end. The model powering the agent doesn't have to be Gemini: our comparison Google Gemini vs ChatGPT vs Claude: which for what use? helps you decide based on the use case, and the monthly comparison of the best LLMs tracks the state of the art. On the deployment side, a well-hosted web page or kiosk is enough to prototype — Hostinger does the job just fine at this stage.

Google's pace is an argument in itself: Live on September 15, Live Avatar on the 24th, TTS the same week. Enterprise building blocks almost always trickle down eventually. Preparing the agent layer now means being ready on the day the face becomes just a configuration option.


❌ Common Mistakes

Mistake 1: Looking for Live Avatar in the Gemini app

The consumer app does not expose the visual persona as of September 24, 2026 — GA is limited to Gemini Enterprise on US and EU endpoints. The solution: go through Gemini Enterprise if your organization has the budget, or work on the audio building block via the API while waiting for broader availability.

Mistake 2: Looking for an "avatar" model ID and its price on the API

The list of live models only shows gemini-3.8-live and gemini-3.8-live-extended-thinking, with no avatar line or public pricing. Looking for "per minute of video" pricing is a dead end. The solution: enterprise contracts and, for freelancers, positioning yourself as an integrator.

Mistake 3: Wanting to clone your own face right away

Technically, a reference photo and an audio sample are enough. Commercially, it's locked behind the Google Cloud sales allowlist. The solution: start with the library of pre-designed avatars, and prepare your assets (clean photo, voice sample) for the day you enter the program.

Mistake 4: Deploying a synthetic face without disclosing it

An avatar that looks like a human, speaks like a human and responds like a human must be disclosed as artificial — a matter of trust as much as compliance. The solution: visible mention in the interface, SynthID kept on the streams, and never a "pretending" posture.


❓ Frequently Asked Questions

Is Live Avatar available in the public Gemini app?

No. As of September 24, 2026, Live Avatar is generally available only in Gemini Enterprise, on US and EU endpoints. The consumer app does not have access to the visual persona, and no public availability date has been announced. Individual developers have no more visibility via the standard API.

How much does Live Avatar cost?

No public pricing has been disclosed: access goes through Gemini Enterprise contracts negotiated with Google. For the audio-only component, gemini-3.8-live is billed at approximately $1.38 per hour (September 2026, verify on ai.google.dev), but no avatar pricing line exists on the standard API to date.

Can you create an avatar from your own photo?

Yes, technically: a reference photo and an audio sample are enough to generate a custom avatar. But the option is strictly reserved for allowlisted enterprise customers, via Google Cloud sales. All other accounts must use the library of pre-built avatars provided by Google, with no facial customization.

Which languages are supported?

97 languages, with native voice-to-face synchronization and mid-conversation language switching. Google claims no degradation of video fidelity and no "visual drift" during these changes, according to the official announcement of September 24, 2026. This native multilingual coverage is presented as the differentiator compared to traditional avatar pipelines.

What's the difference between Gemini 3.8 Live and Live Avatar?

Gemini 3.8 Live, launched on September 15, 2026, is the real-time voice conversation layer, without a speech-to-text pipeline. Live Avatar, announced on September 24, adds a visual persona generated directly by the model. The former is accessible via the Gemini API, the latter only via Gemini Enterprise.

What is the status of Gemini 3.8 Live Extended Thinking?

Private preview. The variant with extended background reasoning is not in GA as of September 24, 2026 and remains behind a waitlist controlled by Google. Interested companies must go through the vendor's sales channels, as with access to custom avatars.


✅ Conclusion

Live Avatar takes the AI assistant from a voice in the void to a visual interlocutor generated by the model itself — a true architectural breakthrough, still behind the enterprise door. While we wait for the wider rollout, build your agent: start with our analysis of Gemini 3.8 Live at $1.38 per hour, then pick your agent's brain in the comparison of the best LLMs for agents.