AI News Feed
Market watch
Large Language Models

Google launches Gemini 3.8 Live voice models with real-time reasoning

Google launched Gemini 3.8 Live and Extended Thinking, voice models with near-real-time reasoning and background tool calls.

In a blog post, Google said the models are designed to run tool and application programming interface calls in the background while maintaining a natural conversational pace. The setup allows an AI agent to keep talking with a user while working on a task just assigned, which Google says creates a more natural, human-like calling experience without constant interruptions when the agent browses the internet for answers.

Google claims top-tier benchmark results for the models. Gemini 3.8 Live Extended Thinking scored 82.6 on the Artificial Analysis Speech to Speech Quality Index, surpassing GPT-Live-1-Astra and Grok Voice Think Fast 2.0, according to Google. Gemini 3.8 Live ranked second on the Speech Agent Arena benchmark and first on ServiceNow’s EVA-Bench. Gemini 3.8 Live Extended Thinking also scored 68.6% on T-Voice, 35.1% on T-Voice-banking and 97.7% on Big Bench Audio, Google said.

The models include automatic language detection and can switch languages mid-conversation. Google said they understand and generate speech in 97 languages, support near-real-time visual grounding, and can use early verbal cues such as “let me check that” to acknowledge a user’s prompt in a more natural way.

Gemini 3.8 Live is available now through the Gemini API and Google AI Studio, and as an enterprise private preview in Gemini Enterprise and Search Live, Google said. Gemini 3.8 Live Extended Thinking is available through the same channels and in Google Workspace through Docs, Gmail and Keep for subscribers, as well as in Gemini Live applications. Developers can integrate the models through partner platforms including Vercel, Agora, LiveKit, Pipecat, Fishjam and Vision Agents, according to Google. Audio files generated by the models carry an invisible SynthID watermark that can be used to detect misinformation.

On pricing, Google said the standard Gemini 3.8 Live model costs $0.005 per minute for audio inputs and $0.018 per minute for outputs. The Extended Thinking model also charges for reasoning tokens and for additional inputs such as video and documents.

Editor's Summary

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice models that aim to reduce latency in voice-based AI agents by combining near-real-time reasoning with background tool calls. The company is offering the models through API, AI Studio, enterprise preview and Workspace channels, with SynthID watermarks on generated audio. Pricing for the standard model starts at $0.005 per minute for audio input and $0.018 per minute for output, while Extended Thinking adds charges for reasoning tokens and additional input types.