Google released Gemini 3.8 Live and 3.8 Live Extended Thinking voice models

Google released two native speech-to-speech models on September 15, 2026: Gemini 3.8 Live, for fast, fluid conversation with real-time visual grounding, and Gemini 3.8 Live Extended Thinking, built for multi-step reasoning that happens while the model keeps speaking. Both detect and switch among 97 languages mid-conversation and can run tools and API calls in the background without breaking the dialogue. The same day Google released Gemini 3.5 Transcribe, a speech-to-text model covering 85-plus languages with a reported average word error rate of 4.0 percent streaming and 2.6 percent non-streaming.

Google reports that the models rank first on the Artificial Analysis Speech to Speech Quality Index with a score of 82.6, score 68.6 percent on tau-Voice agentic tasks and 35.1 percent on Sierra's tau-Voice-banking benchmark, and reach 97.7 percent on Big Bench Audio. Developer pricing for the Live models is listed at $0.005 per minute of audio input and $0.018 per minute of audio output. All generated audio carries SynthID watermarking.

The rollout spans the Gemini API and Google AI Studio, a private preview in Gemini Enterprise, Search Live, the Gemini app, and Google Workspace, where Extended Thinking powers features in Docs, Gmail and Keep. It arrives two months after OpenAI's GPT-Live, and it shows the frontier voice race moving from making models sound natural to making them do agentic work mid-conversation.

The benchmark figures come from Google's own announcement, and several are on leaderboards where rankings shift quickly. A tau-Voice-banking score of 35.1 percent is also a reminder that voice agents still fail most multi-step banking tasks in that test, even at the top of the field.