Google's Gemini 3.8 Live models top speech benchmarks

Google's new Gemini 3.8 Live and Extended Thinking audio models launched September 15. Extended Thinking leads the Artificial Analysis index at 82.6.

16/09/2026 06:119 min read

On September 15, Google unveiled Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as its most capable audio models to date.

These models can converse, reason, and execute tasks. Independent analysts have ranked them: the Extended Thinking variant leads Artificial Analysis’ Speech-to-Speech Index with a score of 82.6, surpassing GPT-Live-1 and Grok Voice.

Say “hi” to our most advanced audio models from @GoogleDeepMind yet built for natural, production-ready voice applications.

đŸ”· Gemini 3.8 Live
đŸ”· Gemini 3.8 Live Extended Thinking

With these models, you can speak naturally, collaborate easily, and tackle complex tasks using
 pic.twitter.com/TCcuANcIWl

— Google (@Google) September 15, 2026

Benchmark Rankings for Google's Latest Conversational AI

Artificial Analysis’ Index combines scores for speech reasoning, agentic performance, arena preference, and task success rate.

When tested with high reasoning effort, Gemini 3.8 Live Extended Thinking took the top spot. GPT-Live-1 Astra came next with 81.5, and Grok Voice Think Fast 2.0 High scored 81.3.

The regular Gemini 3.8 Live ranked fifth at 76.0. Both new models outperformed Gemini 3.1 Flash Live High, which had 71.5.

The difference in agentic capability is larger. Extended Thinking achieved 68.6% on the Tau Voice benchmark, compared to 37.7% for the earlier model.

In speech reasoning, Extended Thinking posted 97.7%, just ahead of Grok Voice’s 97.2%. But it fell short of Qwen Audio 3.0 Realtime Plus, which reached 99.2%.

Older Gemini Model Still Preferred by Human Testers

Pricing shows a clear difference. The standard model costs $0.84 per hour of input audio, about half the predecessor's $1.75 and the cheapest in the Index.

Extended Thinking is priced at $3.50 per hour, lower than GPT-Live-1 Sol's $4.47 and Grok Voice Think Fast 2.0 High's $4.80.

Latency also improved. The average time to first audio dropped to 1.18 seconds from 2.99 seconds for the older model.

But human listeners are not entirely convinced. In blind Speech Agent Arena conversations, Gemini 3.1 Flash Live continues to lead in preference with an Elo rating of 1096.

Gemini 3.8 Live ranks second with 1083, while the Extended Thinking version lags at 990, even though it completes 89.1% of tasks.

Voice AI evaluation now uses two distinct metrics heading in opposite directions. Benchmarks favor the model with strongest reasoning, while preference scores favor the best conversationalist. Google holds the lead in both categories, albeit with different models.

Share to

Disclaimer: this article comes from third-party media and is provided for reference only. It does not constitute investment advice. Crypto and other financial products carry significant price volatility risk, so please make your own decisions carefully.

Related articles