Skip to content
KoishiAI
ไทย
← Back to all articles

Gemini 3.8 Live vs Extended Thinking: Benchmarks & API

Explore Gemini 3.8 Live vs Extended Thinking: benchmark scores, language support, API access for native speech-to-speech enterprise voice agents.

AI-drafted from cited sources, fact-checked and reviewed by a human editor. How we work · Standards · Report an error
A digitally rendered abstract image showcasing a futuristic eye with complex network patterns.
Photo by Merlin Lightpainting on Pexels

TL;DR: Google DeepMind’s Gemini 3.8 Live Extended Thinking tops Artificial Analysis’ index with an 82.6 score by replacing fragmented speech pipelines with native audio processing. This unified approach enables real-time reasoning and visual grounding, setting a new standard for latency-free enterprise voice agents.

Key facts

  • Google DeepMind announced Gemini 3.8 Live and Extended Thinking on September 15, 2026, as native speech-to-speech models replacing cascaded pipelines.
  • Gemini 3.8 Live supports automatic detection for 97 languages and features visual grounding capabilities for near real-time image or video processing.
  • The Gemini 3.8 Live Extended Thinking model holds the top position on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6.
  • Extended Thinking achieved 68.6% on the τ-Voice benchmark, 35.1% on Sierra’s τ-Voice-banking benchmark, and 97.7% on Big Bench Audio.
  • Both models are available via the Gemini Live API and Google AI Studio as hosted services only, with no self-hosted or open-weight versions offered.

Replacing Cascaded Pipelines with Native Speech Models

Google DeepMind has released two new voice-focused AI models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed to replace traditional speech processing systems for enterprise applications [1][2]. Announced on September 15, 2026, these native speech-to-speech models aim to eliminate the latency and fragmentation of older “cascaded” pipelines by integrating reasoning, tool use, and audio generation into a single continuous flow [4].

The release marks a significant shift in how AI voice agents are built. Instead of relying on separate components for speech recognition, language processing, and text-to-speech synthesis, these models process input and output directly as audio streams [1][2]. This architecture allows the AI to maintain natural conversational rhythm while simultaneously executing background tasks or analyzing visual data without breaking the user’s experience [4].

Model Capabilities and Performance

The two new models serve different tiers of complexity within the Gemini Audio family, which was previously expanded with the Gemini 3.5 Transcribe model last month [4].

Gemini 3.8 Live: Scale and Efficiency

Gemini 3.8 Live is optimized for high-volume, cost-efficient applications where speed and fluidity are prioritized over deep analytical reasoning [1][2]. It features visual grounding capabilities, allowing it to process images or video inputs in near real-time to enrich the conversation with contextual details [1][2][4].

The model supports automatic detection for 97 languages, making it suitable for global customer service and multilingual applications [3]. A key technical feature is its asynchronous function calling; this allows the AI to execute background tasks, such as querying databases or scheduling meetings, while continuing to stream audio responses to the user without interruption [3][4].

In human preference evaluations on the Speech Agent Arena, Gemini 3.8 Live secured second place against competing voice agents [1][2][4]. This indicates strong performance in naturalness and responsiveness, though it trades some analytical depth for speed.

Gemini 3.8 Live Extended Thinking: Complex Reasoning

For tasks requiring multi-step planning and higher intelligence, Google introduced Gemini 3.8 Live Extended Thinking [1][2][4]. This model is designed to handle complex workflows where accuracy and logical consistency are critical [1][2].

The model currently holds the top position on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6 [1][2][4]. Its performance in agentic tasks—where AI agents must complete multi-step goals—is notably high: it achieved 68.6% on the τ-Voice benchmark and 35.1% on Sierra’s τ-Voice-banking benchmark [1][2][4]. It also scored 97.7% on Big Bench Audio, a standard test for audio understanding capabilities [1][2][4].

Google reports that on ServiceNow’s EVA-Bench, which measures performance in complex enterprise workflows, these models push the “Pareto Frontier” by balancing high task accuracy with natural conversational quality [1][2][4]. This suggests they are particularly effective for sophisticated customer support or technical assistance roles where both correctness and user experience matter.

Availability and Ecosystem Integration

Both Gemini 3.8 Live models are available through the Gemini Live API and Google AI Studio [4]. However, Google clarified that these are hosted services only; there is no self-hosted or open-weight version of these specific models available for developers to run on their own hardware [4].

This hosting model aligns with Google’s broader strategy for enterprise integration. The models are designed to work seamlessly within the Gemini app, Google Workspace, and Search, providing a unified infrastructure for voice-enabled productivity tools [1][2].

Context Within the Gemini 3.8 Lineup

The Live models join a expanding family of Gemini 3.8 variants released in September 2026. Earlier in the month, Google launched Gemini 3.8 Flash and Gemini 3.8 Flash Cyber [5][6]. These earlier releases focused on software engineering and cybersecurity tasks respectively, targeting developers rather than voice interaction [7].

While the Flash models prioritize code generation and security analysis, the Live models address the growing demand for real-time voice agents in customer service, telephony, and interactive assistants [4]. This diversification allows enterprises to select models based on their specific needs: speed and cost for routine interactions (Live), deep reasoning for complex problem-solving (Extended Thinking), or code/analysis for technical workflows (Flash).

The introduction of native speech-to-speech capabilities represents a move away from the modular AI architectures that dominated previous years. By unifying perception, reasoning, and generation into a single model, Google DeepMind aims to reduce latency and improve the naturalness of AI voice interactions [4]. As enterprises continue to adopt voice agents for customer support and internal operations, these benchmarks provide a clearer picture of where current technology stands in terms of both capability and reliability.

Sources

  1. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (deepmind.google) — 2026-09-15
  2. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (blog.google) — 2026-09-15
  3. Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents (www.marktechpost.com) — 2026-09-15
  4. Google DeepMind (@GoogleDeepMind) on X (x.com) — 2026-09-15
  5. Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (blog.google) — 2026-09-02
  6. Gemini 3.8 Flash Cyber (radar.mahdi.uk) — 2026-09-02
  7. Google DeepMind Launches Gemini 3.8 Flash Cyber? How AI Defense Works (www.opoinstall.com) — 2026-09-03

Frequently asked questions

What is the difference between Gemini 3.8 Live and Extended Thinking?
Gemini 3.8 Live is optimized for high-volume applications prioritizing speed and cost-efficiency, while Gemini 3.8 Live Extended Thinking is designed for complex workflows requiring multi-step planning and higher logical consistency. The standard Live model trades some analytical depth for fluidity, whereas the Extended Thinking version holds the top position on Artificial Analysis’ Speech to Speech Quality Index.
Can I download and host Gemini 3.8 Live locally?
These models are available exclusively through the Gemini Live API and Google AI Studio as hosted services only. Google has clarified that there is no self-hosted or open-weight version of these specific native speech-to-speech models for developers to run on their own hardware.
How many languages does Gemini 3.8 Live support?
Gemini 3.8 Live supports automatic detection for 97 languages, making it suitable for global customer service and multilingual applications. This broad language support allows the model to handle diverse international voice interactions without requiring separate translation pipelines.
How does Gemini 3.8 Live improve upon previous voice AI architectures?
These models replace traditional cascaded pipelines by integrating speech recognition, reasoning, and audio generation into a single continuous flow. This native architecture eliminates the latency and fragmentation of older systems that relied on separate components for each step.
What are the benchmark scores for Gemini 3.8 Live Extended Thinking?
The Extended Thinking model achieved 68.6% on the τ-Voice benchmark, 35.1% on Sierra’s τ-Voice-banking benchmark, and 97.7% on Big Bench Audio. These scores indicate strong performance in agentic tasks and audio understanding capabilities compared to other voice agents.