Speech to Speech AI Model & Provider Leaderboard

Analysis and comparison of Speech to Speech models & API providers. Artificial Analysis has analyzed speech to speech models and hosting providers across different characteristics including their reasoning quality, conversational dynamics, generation time and price.

For further details, see the methodology page.

Highlights

Weighted average of Speech Reasoning, Agentic Performance, Arena Preference, and Task Success Rate results · Higher is better
Time to first audio (seconds) on Big Bench Audio · Lower is better
Cost to complete fixed 40-question Big Bench Audio subset based on length of input audio, normalized to hourly basis · Lower is better

Artificial Analysis Speech to Speech IndexUpdated

Artificial Analysis Speech to Speech Index

Weighted average of Speech Reasoning, Agentic Performance, Arena Preference, and Task Success Rate results · Only models with all data available shown · Higher is better

Speech Reasoning

Speech Reasoning (Big Bench Audio)

Speech reasoning: based on the Artificial Analysis Big Bench Audio dataset · Higher is better

Agentic Performance

Agentic Performance (𝜏-Voice)

Proportion of replica customer service scenarios resolved while acting as a customer support agent, based on the 𝜏-Voice benchmark · Higher is better · Only full duplex models

Note: Following models based on 1 trial: GPT-Realtime-2 (Minimal), OpenAI; Following models based on 2 trials: GPT-Realtime-2.1 High, OpenAI

Agentic Performance (𝜏-Voice) by Domain

Proportion of replica customer service scenarios resolved by domain, based on the 𝜏-Voice benchmark · Higher is better · Only full duplex models

Note: Following models based on 1 trial: GPT-Realtime-2 (Minimal), OpenAI; Following models based on 2 trials: GPT-Realtime-2.1 High, OpenAI

Speech Agent ArenaNew

Arena Preference Elo

Preference Elo measures human participants' overall preference after they complete the same scenario with two unidentified Speech to Speech models, such as booking a dental appointment · Only includes models with tool calling · Higher is better

Task Success Rate

Proportion of eligible conversations with the correct final task-completing tool call or calls · Higher is better

Conversational Dynamics

Conversational Dynamics (Full Duplex Bench subset)

Weighted average of pause handling, turn-taking, interruption handling, and backchannel handling from Full Duplex Bench v1 and v1.5 · Higher is better

Conversational Dynamics (Full Duplex Bench subset) - Category Breakdown

Individual category scores ordered by conversational dynamics score · Based on subset of Full Duplex Bench v1 and v1.5

API Benchmarks

Cost per Hour of Input Audio (Big Bench Audio subset)

Cost to complete fixed 40-question Big Bench Audio subset based on length of input audio, normalized to hourly basis · Lower is better

Speech Reasoning (Big Bench Audio) vs. Cost per Hour of Input Audio

Big Bench Audio score vs. cost per hour of input audio measured on fixed 40-question subset · Higher speech reasoning and lower cost are better
Most attractive quadrant

Time to First Audio

Average time to first audio (seconds) on Big Bench Audio dataset · Lower is better

Speech Reasoning (Big Bench Audio) vs. Speed

Big Bench Audio score vs. time to first audio · Higher speech reasoning and lower latency are better
Most attractive quadrant

Summary of Key Metrics & Further Information

Alibaba Cloud logoAlibaba Cloud
Qwen Audio 3.0 Realtime Plus
66.8%
99%
98.4%
54.6%
699
77.8%
1.54
4.42
0.03
0.18
Alibaba Cloud logoAlibaba Cloud
Qwen3.5 Omni Plus Realtime
-
99%
-
-
799
62.0%
2.64
0.00
0.16
0.82
Alibaba Cloud logoAlibaba Cloud
Qwen Audio 3.0 Realtime Flash
64.2%
96%
96.9%
35.9%
752
81.7%
1.55
4.77
-
-
Alibaba Cloud logoAlibaba Cloud
Qwen3.5 Omni Flash Realtime
-
59%
-
-
797
29.1%
0.79
0.00
0.16
0.82
Alibaba Cloud logoAlibaba Cloud
Qwen3 Omni Flash
-
59%
72.7%
-
-
-
4.82
1.77
0.61
1.36
Alibaba Cloud logoAlibaba Cloud
Qwen3 Omni Realtime
-
57%
-
-
-
-
0.88
2.26
0.65
0.99
Amazon Bedrock logoAmazon Bedrock
Nova 2.0 Sonic (Mar 2026)
-
88%
-
-
917
57.1%
1.14
-
0.27
1.08
Boson AI logoBoson AI
Higgs Realtime
-
69%
92.8%
18.6%
-
-
1.47
-
-
-
Cofe AI logoCofe AI
FLM-Audio
-
16%
62.0%
-
-
-
-
-
-
-
Deepslate logoDeepslate
Deepslate Opal
-
85%
85.7%
17.5%
-
-
0.44
6.48
-
-
Google logoGoogle
Gemini 3.1 Flash Live High
71.5%
97%
74.3%
37.7%
1014
71.8%
2.99
1.75
0.35
1.38
Google logoGoogle
Gemini 2.5 Flash Native Audio Dialog Thinking
-
91%
-
-
-
-
3.87
-
0.35
1.38
Google logoGoogle
Gemini 3.1 Flash Live Minimal
63.9%
71%
72.3%
26.2%
1046
74.6%
0.96
1.50
0.35
1.38
Google logoGoogle
Gemini 2.5 Flash Native Audio Dialog
-
69%
-
-
-
-
0.63
1.42
0.35
1.38
Google logoGoogle
Gemini 2.5 Flash Native Audio Preview (Dec 2025)
-
-
44.0%
22.8%
-
-
-
-
-
-
Google logoGoogle
Gemini 2.5 Flash Native Audio Preview (Sep 2025)
-
-
30.3%
-
-
-
-
-
-
-
Kyutai logoKyutai
Moshi
-
4%
61.0%
-
-
-
-
-
-
-
NVIDIA logoNVIDIA
Nemotron Voicechat
-
27%
52.9%
-
-
-
-
-
-
-
NVIDIA logoNVIDIA
PersonaPlex
-
19%
91.0%
-
-
-
-
-
-
-
OpenAI logoOpenAI
GPT-Realtime-2 (High)
73.6%
97%
95.3%
39.8%
914
89.8%
1.14
4.14
1.15
4.61
OpenAI logoOpenAI
GPT-Realtime-2.1 High
73.9%
96%
95.7%
45.7%
892
91.5%
1.21
10.75
-
-
OpenAI logoOpenAI
GPT-Realtime-2 (Medium)
-
93%
95.2%
37.4%
-
-
1.22
3.97
1.15
4.61
OpenAI logoOpenAI
GPT-Realtime-2.1 Minimal
70.3%
87%
92.7%
38.0%
896
89.4%
0.97
11.31
-
-
OpenAI logoOpenAI
GPT Realtime (Aug '25)
68.5%
83%
93.9%
30.4%
944
89.4%
0.98
11.08
1.15
4.61
OpenAI logoOpenAI
GPT-Realtime-1.5
70.3%
81%
95.7%
38.8%
1000
85.1%
0.81
11.44
1.15
4.61
OpenAI logoOpenAI
GPT-Realtime-2.1 Mini High
-
75%
91.7%
29.4%
-
-
4.28
3.45
-
-
OpenAI logoOpenAI
GPT-Realtime-2 (Minimal)
62.7%
72%
96.1%
30.8%
881
84.7%
1.12
3.07
1.15
4.61
OpenAI logoOpenAI
GPT-4o mini Realtime (Dec 2024)
-
69%
-
-
-
-
1.27
5.75
0.36
1.44
OpenAI logoOpenAI
GPT Realtime Mini (Oct '25)
56.8%
64%
95.7%
15.1%
912
79.6%
0.81
3.04
0.36
1.44
OpenAI logoOpenAI
GPT-Realtime-2.1 Mini Minimal
52.8%
63%
91.8%
22.5%
792
76.7%
0.85
4.60
-
-
OpenAI logoOpenAI
GPT-4o audio chatcompletions
-
54%
-
-
-
-
3.38
0.00
3.60
14.40
OpenAI logoOpenAI
GPT-4o Realtime (Dec 2024)
-
-
89.8%
27.9%
-
-
1.49
2.04
1.44
5.76
SpaceXAI logoSpaceXAI
Grok Voice Think Fast 2.0 High
79.0%
97%
95.1%
56.5%
908
94.7%
0.70
4.80
-
-
SpaceXAI logoSpaceXAI
Grok Voice Think Fast 1.0
72.3%
97%
77.8%
52.1%
839
80.7%
1.25
3.00
-
-
SpaceXAI logoSpaceXAI
Grok Voice Fast 1.0
-
93%
71.6%
27.4%
-
-
0.78
3.00
3.00
3.00
StepFun logoStepFun
Step-Audio R1.1 (Realtime)
-
98%
-
-
-
-
1.53
0.00
0.06
1.69
VITA logoVITA
Freeze-Omni
-
33%
58.7%
-
-
-
-
-
-
-

Frequently Asked Questions