Connect with us

Daily News

Google expands Gemini world with 3.8 Live and 3.8 Extended Thinking

Published

on

Expanding the Gemini world, Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These are introduced as a part of the advanced live dialogue models. These models as mentioned by Google will deliver the building blocks for reliable, production-ready voice agents. 

Gemini 3.8 Live is built for scale and cost efficiency. It’s meant to power high-volume, production-ready voice agents. They need to combine conversational intelligence, fluid dialogue and visual grounding without racking up heavy compute costs.

Gemini 3.8 Live Extended Thinking, on the other hand, is aimed at more complex tasks. It adds deeper multi-step reasoning on top of the same real-time voice capabilities. 

Google has published demos showing the models in action like how they guide employee onboarding using visual context, playing chess in real time, turning hand-drawn sketches into functional React components through live voice feedback and more. 

The demo video also shows how the extended version of Gemini coordinates multi-step restaurant bookings in the background during a natural conversation. It also builds out business plans and marketing toolkits on the fly, entirely through speech.

The Live models are also making their way into Google’s everyday products. In Google Workspace, Extended Thinking powers voice-driven navigation and real-time drafting across Docs Live, Gmail Live and Keep Live. In Search, it enables step-by-step, real-time voice troubleshooting through Search Live.

This will free the developers to not build real-time voice infrastructure from scratch. Google is leaning on its Gemini Live API and a network of platform partners — including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents — which handle the underlying media streaming so developers can focus on the actual user experience.

Enterprise partners including Salesforce, Genspark and Lumeris have also been testing the models. These partners shared their feedback with Google on latency, conversational fluidity and tool-calling performance.

In line with its existing responsible-AI practices, Google confirmed that all audio generated by these models is watermarked using SynthID. This is an imperceptible marker embedded directly into the audio output, intended to keep AI-generated content identifiable and help curb misinformation.