Gemini 3.8 Live vs Extended Thinking: Benchmarks & API
Explore Gemini 3.8 Live vs Extended Thinking: benchmark scores, language support, API access for native speech-to-speech enterprise voice agents.
TL;DR: Google DeepMind’s Gemini 3.8 Live Extended Thinking tops Artificial Analysis’ index with an 82.6 score by replacing fragmented speech pipelines with native audio processing. This unified approach enables real-time reasoning and visual grounding, setting a new standard for latency-free enterprise voice agents.
Key facts
- Google DeepMind announced Gemini 3.8 Live and Extended Thinking on September 15, 2026, as native speech-to-speech models replacing cascaded pipelines.
- Gemini 3.8 Live supports automatic detection for 97 languages and features visual grounding capabilities for near real-time image or video processing.
- The Gemini 3.8 Live Extended Thinking model holds the top position on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6.
- Extended Thinking achieved 68.6% on the τ-Voice benchmark, 35.1% on Sierra’s τ-Voice-banking benchmark, and 97.7% on Big Bench Audio.
- Both models are available via the Gemini Live API and Google AI Studio as hosted services only, with no self-hosted or open-weight versions offered.
Replacing Cascaded Pipelines with Native Speech Models
Google DeepMind has released two new voice-focused AI models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed to replace traditional speech processing systems for enterprise applications [1][2]. Announced on September 15, 2026, these native speech-to-speech models aim to eliminate the latency and fragmentation of older “cascaded” pipelines by integrating reasoning, tool use, and audio generation into a single continuous flow [4].
The release marks a significant shift in how AI voice agents are built. Instead of relying on separate components for speech recognition, language processing, and text-to-speech synthesis, these models process input and output directly as audio streams [1][2]. This architecture allows the AI to maintain natural conversational rhythm while simultaneously executing background tasks or analyzing visual data without breaking the user’s experience [4].
Model Capabilities and Performance
The two new models serve different tiers of complexity within the Gemini Audio family, which was previously expanded with the Gemini 3.5 Transcribe model last month [4].
Gemini 3.8 Live: Scale and Efficiency
Gemini 3.8 Live is optimized for high-volume, cost-efficient applications where speed and fluidity are prioritized over deep analytical reasoning [1][2]. It features visual grounding capabilities, allowing it to process images or video inputs in near real-time to enrich the conversation with contextual details [1][2][4].
The model supports automatic detection for 97 languages, making it suitable for global customer service and multilingual applications [3]. A key technical feature is its asynchronous function calling; this allows the AI to execute background tasks, such as querying databases or scheduling meetings, while continuing to stream audio responses to the user without interruption [3][4].
In human preference evaluations on the Speech Agent Arena, Gemini 3.8 Live secured second place against competing voice agents [1][2][4]. This indicates strong performance in naturalness and responsiveness, though it trades some analytical depth for speed.
Gemini 3.8 Live Extended Thinking: Complex Reasoning
For tasks requiring multi-step planning and higher intelligence, Google introduced Gemini 3.8 Live Extended Thinking [1][2][4]. This model is designed to handle complex workflows where accuracy and logical consistency are critical [1][2].
The model currently holds the top position on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6 [1][2][4]. Its performance in agentic tasks—where AI agents must complete multi-step goals—is notably high: it achieved 68.6% on the τ-Voice benchmark and 35.1% on Sierra’s τ-Voice-banking benchmark [1][2][4]. It also scored 97.7% on Big Bench Audio, a standard test for audio understanding capabilities [1][2][4].
Google reports that on ServiceNow’s EVA-Bench, which measures performance in complex enterprise workflows, these models push the “Pareto Frontier” by balancing high task accuracy with natural conversational quality [1][2][4]. This suggests they are particularly effective for sophisticated customer support or technical assistance roles where both correctness and user experience matter.
Availability and Ecosystem Integration
Both Gemini 3.8 Live models are available through the Gemini Live API and Google AI Studio [4]. However, Google clarified that these are hosted services only; there is no self-hosted or open-weight version of these specific models available for developers to run on their own hardware [4].
This hosting model aligns with Google’s broader strategy for enterprise integration. The models are designed to work seamlessly within the Gemini app, Google Workspace, and Search, providing a unified infrastructure for voice-enabled productivity tools [1][2].
Context Within the Gemini 3.8 Lineup
The Live models join a expanding family of Gemini 3.8 variants released in September 2026. Earlier in the month, Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber [5][6]. These earlier releases focused on software engineering and cybersecurity tasks respectively, targeting developers rather than voice interaction [7].
While the Flash models prioritize code generation and security analysis, the Live models address the growing demand for real-time voice agents in customer service, telephony, and interactive assistants [4]. This diversification allows enterprises to select models based on their specific needs: speed and cost for routine interactions (Live), deep reasoning for complex problem-solving (Extended Thinking), or code/analysis for technical workflows (Flash).
The introduction of native speech-to-speech capabilities represents a move away from the modular AI architectures that dominated previous years. By unifying perception, reasoning, and generation into a single model, Google DeepMind aims to reduce latency and improve the naturalness of AI voice interactions [4]. As enterprises continue to adopt voice agents for customer support and internal operations, these benchmarks provide a clearer picture of where current technology stands in terms of both capability and reliability.
Sources
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (deepmind.google) — 2026-09-15
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (blog.google) — 2026-09-15
- Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents (www.marktechpost.com) — 2026-09-15
- Google DeepMind (@GoogleDeepMind) on X (x.com) — 2026-09-15
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (blog.google) — 2026-09-02
- Gemini 3.8 Flash Cyber (radar.mahdi.uk) — 2026-09-02
- Google DeepMind Launches Gemini 3.8 Flash Cyber? How AI Defense Works (www.opoinstall.com) — 2026-09-03