
Gemini 3.8 Live, Google's new family of real-time voice models, was announced on 15 September 2026. There are two models: Gemini 3.8 Live for fast, low-cost conversation, and Gemini 3.8 Live Extended Thinking for harder tasks that need step-by-step reasoning. Both can keep talking while they look something up or carry out a task in the background, and the standard model can move between 97 languages in the middle of a conversation. They began rolling out the same day to developers, to Google Search and the Gemini app, and to voice features in Google Workspace.
What happened
Google described the release in its post introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, written by members of the Gemini Audio team. The company calls them "our most advanced live dialogue models yet".
A live model is different from a chatbot that reads its answers aloud. It takes in speech, and in this case images and video, as a continuous stream and answers in speech, so the exchange works like a phone call. The two new models split the job:
- Gemini 3.8 Live is built for scale and cost efficiency. It processes visual input in near real time, detects and switches between 97 supported languages mid-conversation, and runs tools and calls to other systems in the background while it carries on talking.
- Gemini 3.8 Live Extended Thinking is built for complex tasks. Google says it reasons and speaks at the same time, acknowledges a request with a short phrase such as "Let me check that…" and then narrates its progress through multi-step work.
Google's demonstrations include guiding a new employee through onboarding using visual context, coordinating a multi-step booking, and troubleshooting a problem step by step inside Search.
Key details
| Detail | Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking |
|---|---|---|
| Designed for | High-volume, cost-sensitive conversation | Complex, multi-step tasks |
| Developers | Gemini API and Google AI Studio | Gemini API and Google AI Studio |
| Enterprises | Private preview in Gemini Enterprise | Private preview in Gemini Enterprise. Coming soon for Google Workspace business customers. |
| Everyone | Search Live | Gemini Live. Docs for Google AI Pro and Ultra subscribers. Gmail and Keep for all Google AI subscribers. |
| Results Google reports | Second place in the Speech Agent Arena, a ranking based on user preference | 82.6 on the Artificial Analysis Speech to Speech Quality Index and 97.7 per cent on Big Bench Audio |
Google's model card for Gemini 3.8 Audio adds the technical limits: the Live models accept audio, images, video and text with a context window of up to 128K tokens, which is the amount of conversation and material they can keep in mind at once. The announcement adds that all audio generated by Google's AI products is watermarked with SynthID, an imperceptible watermark meant to keep AI-generated audio detectable. The announcement gives no prices, describing the models only as cost-effective. Thurrott's report on the announcement confirms the same availability for developers, enterprises and Search Live.
Why it matters
The awkward pause is the target. Google's pitch is that the silence while a voice assistant works on something goes away. A model that says it is checking, keeps the caller informed and comes back with a result behaves much more like a person on a service desk. That is what makes a voice agent tolerable for a real booking or enquiry.
Voice is now a way to do things, not only to ask things. Running tools in the background means the model can look up an order, check a calendar or fill in a form while the conversation continues. Google names Salesforce among the companies working with the models, and lists developer platforms that handle the real-time audio plumbing so that a software team can concentrate on the experience.
The scores show progress and limits together. Google reports that Extended Thinking completes 68.6 per cent of tasks on a voice agent benchmark called τ-Voice and 35.1 per cent on a banking version of it built by Sierra. Google presents both as the top results available. Read the other way, about two in three tasks on that banking benchmark are still not completed. The model card also lists hallucinations, meaning confident wrong answers, and occasional slowness or timeouts among known limitations.
Language switching has everyday uses. A model that follows a speaker from one language into another and back suits a multilingual customer base. AI that works with speech is also used in language learning: PractiseNow, for example, offers speaking practice for English tests such as IELTS, PTE and OET with AI marking of spoken answers.
What this means for businesses
There are two ways this reaches a small or medium business. The first is as a user. If your team works in Google Workspace, voice features in Gmail, Docs and Keep are arriving for people with Google AI subscriptions, and Google says Workspace business customers are coming soon. Dictating a reply or talking through a document is convenient, and it also means work content is being spoken aloud in shared offices and processed by an AI feature, so the usual rules about what goes into AI tools apply.
The second is as something your customers talk to. A clinic's front desk, a restaurant taking bookings or a trades business that misses calls while on site are the obvious candidates for a voice agent that answers, takes details and books a time. The technology is now good enough to be worth a trial and not yet good enough to leave unsupervised. Start with a narrow task, such as taking a message or booking a standard appointment, and give callers a quick way to reach a person.
Be open about it. The Office of the Australian Information Commissioner says in its guidance on commercially available AI products that public-facing AI tools should be clearly identified as such to the people using them, and that privacy obligations apply to personal information entered into an AI system and to the output it generates. A phone call usually contains a name, a number and a reason for calling, which for a health or legal practice can be sensitive.
A short checklist:
- Check whether voice features are switched on in your Google Workspace and who can use them.
- Pick one narrow phone task to trial, and write down what the agent must never do.
- Tell callers at the start that they are speaking with an automated assistant.
- Make the handover to a person simple and test it.
- Decide where call audio and transcripts are stored, for how long, and who can see them.
- Review a sample of calls every week during the trial.
Comingwave is a technology company that provides business systems and integrations, custom software and managed IT support to small and medium businesses, including hospitality venues that live on bookings. A voice agent is only as useful as the booking system, calendar or CRM it connects to. If you would like help getting those systems ready, request a free first consultation. We reply within one business day.
Key takeaways
- Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026.
- Both can carry out tasks in the background while continuing to talk, and the standard model switches between 97 languages mid-conversation.
- They are available to developers now and are rolling out in Search Live, Gemini Live and Workspace voice features.
- Google's own benchmark figures show strong progress and a meaningful failure rate on complex tasks.
- Businesses trialling voice agents should start narrow, disclose the automation and keep a person within reach.
Frequently asked questions
What is Gemini 3.8 Live?
Gemini 3.8 Live is a real-time voice model from Google, announced on 15 September 2026. It listens and speaks as a continuous conversation, can take in visual input, and can run tasks in the background while it talks.
What is the difference between Gemini 3.8 Live and Extended Thinking?
Gemini 3.8 Live is the faster, lower-cost model for everyday conversation at scale. Extended Thinking is designed for complex tasks: it reasons through several steps while speaking and narrates its progress.
How many languages does Gemini 3.8 Live support?
Google says Gemini 3.8 Live automatically detects and switches between 97 supported languages in the middle of a conversation.
Can a small business use Gemini 3.8 Live to answer phone calls?
Not out of the box. The models are available to developers through the Gemini API, so a phone agent has to be built or bought from a provider that uses them, and then connected to your booking or customer systems.
Is AI-generated speech from Gemini labelled?
Google says all audio generated by its AI products is watermarked with SynthID, an imperceptible marker in the audio. That helps with later detection. It does not tell a caller in the moment, so a spoken disclosure is still needed.