On September 15, 2026, Google officially released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, positioning them as its most advanced real-time conversational models to date, and opened API access two days later. According to the Artificial Analysis speech-to-speech quality index, the Extended Thinking version ranks first worldwide with a score of 82.6%; it reaches 68.6% on the τ-Voice agent task completion benchmark, 35.1% on Sierra's τ-Voice-banking financial scenario benchmark, and 97.7% on the Big Bench Audio reasoning test. This is the first time Google has merged real-time reasoning capability with voice conversation into a single model, and both products are open to developers through the Gemini API and Google AI Studio.
How the Two Models Divide the Technical Work
Rather than launching a single product, Google split this release by scenario: Gemini 3.8 Live targets scaled deployment and cost efficiency, positioned as the workhorse engine for high-concurrency scenarios; the Extended Thinking version targets highly complex tasks and comes equipped with stronger multi-step reasoning. Google's official blog describes the positioning of the two models as "one providing developers and enterprises with a reliable, production-ready foundation for voice agents, and the other making conversations smoother and collaboration more intuitive."
The most critical breakthrough in the Extended Thinking version is "reasoning and speaking in parallel." In tasks that require multi-step background processing, the model does not fall silent; instead it tells the user the current progress through natural-language cues such as "Let me check on that...," while continuing to work through the logic in the background. Google's official blog describes this as "real-time narration of multi-step background task progress without sacrificing conversational fluency."
The highlights of the base Gemini 3.8 Live lean toward breadth: it supports near-real-time visual input, can automatically detect and switch among 97 supported languages during a conversation, and performs background tool and API calls without interrupting the dialogue.
Google has also embedded SynthID invisible watermarks into all AI-generated audio, written directly into the audio output so that AI-generated content can be detected later.
The Many Facets of the Benchmark Numbers
Google chose Artificial Analysis's scores as its headline promotional data: an 82.6% speech-to-speech quality index, currently at the top of the leaderboard, and Artificial Analysis is an independent third-party evaluation body. The τ-Voice score of 68.6% and the Big Bench Audio reasoning score of 97.7% form a complete set of strengths.
In the Speech Agent Arena real-user preference Elo ranking, Gemini 3.8 Live sits second with 1,083 points, behind Gemini 3.1 Flash Live (1,096 points). In human conversation evaluations, the previous-generation version of the older model still beat the new release. OpenAI's GPT-Live-1 (the low-config Sol version) ranks third with 1,053 points.
User preference arenas generally reflect perceived dimensions such as voice naturalness, response pacing, and conversational coherence, and these dimensions do not have a linear relationship with reasoning accuracy. The new model is stronger on reasoning tasks, but if its speech prosody or interruption handling is not natural enough, user experience scores will be discounted accordingly. This shows that Google has currently prioritized "smarter" over "more human," and the two are separable engineering directions.
On the task success rate dimension, Grok Voice Think Fast 2.0 High leads with 94.6%, with Gemini 3.8 Live close behind at 93.2%. No single model ranks first on all three leaderboards, and the voice AI field is still in a phase of multi-dimensional divergence.
Sierra's τ-Voice-banking score of 35.1% specifically tests multi-step business process handling in financial scenarios and is widely recognized in the industry as a difficult benchmark. A score of 35.1% is not low, but it also means there is still a visible engineering gap before that scenario can reach fully intervention-free commercial deployment. Google is directing enterprise customers to an "early access" channel rather than opening it up fully right away.
Industry Impact for Developers and Enterprises
For developers, the substantive value of this release lies in how complete the ecosystem access paths are. Google has completed integrations with mainstream development platforms including Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents, so developers can call the Gemini Live API within familiar toolchains without rebuilding infrastructure. Google also notes that Gemini 3.8 Live "maintains competitive pricing," but did not publish specific figures in the release materials.
On the enterprise side, Google announced partnerships with Salesforce, Genspark, and Lumeris to advance ecosystem building, targeting CRM customer service automation, intelligent assistant platforms, and healthcare scenarios respectively. The Salesforce partnership is especially critical — Salesforce has an enormous enterprise customer base, and if Gemini 3.8 Live can be deeply embedded in its customer service platform, it can reach a vast number of real business scenarios without relying on standalone sales efforts.
On the consumer side, Google has built a complete ladder from free to paid: all users can try the basic voice features in Search Live; Gemini Live users can access the Extended Thinking version; Google AI Pro/Ultra subscribers can use the reasoning version in Workspace Docs; and Gmail and Keep are open to all Google AI subscribers.
Strategic Assessment
This release reveals a clear product bet: in the next round of competition in voice AI, "parallel reasoning" rather than "faster responses" will become the core battleground for differentiation. Once the latency of mainstream voice models all enters an acceptable range, whether a model can complete multi-step complex tasks without interrupting the conversation will become the true dividing line in enterprise customer selection. Extended Thinking's "think while you speak" mechanism is essentially placing an early bet on that judgment.
Whether the τ-Voice-banking score of 35.1% can rise quickly through subsequent iterations will directly determine the pace of commercialization in highly regulated industries such as finance and insurance. Whether Gemini 3.1 Flash Live's suppression of the new model in the user preference arena can be reversed in the coming months is another question — if it can never be reversed, that would indicate a systematic tuning problem in Google's voice perception dimensions rather than a gap in algorithmic capability itself, and the two require completely different remediation paths.
Salesforce integration progress is another milestone. If this partnership achieves scaled deployment before the end of 2026, it will be a landmark event for voice AI entering mainstream channels of traditional enterprise software, and its demonstration effect will far exceed any single technical benchmark score improvement.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接