Speech to Speech AI Model & Provider Leaderboard
Analysis and comparison of Speech to Speech models & API providers. Artificial Analysis has analyzed speech to speech models and hosting providers across different characteristics including their reasoning quality, conversational dynamics, generation time and price.
For further details, see the methodology page.
Highlights
Artificial Analysis Speech to Speech Index
Artificial Analysis Speech to Speech Index
Speech Reasoning
Speech Reasoning (Big Bench Audio)
Agentic Performance
Agentic Performance (𝜏-Voice)
Note: Following models based on 1 trial: GPT-Realtime-2 (Minimal), OpenAI; Following models based on 2 trials: GPT-Realtime-2.1 High, OpenAI, OpenAI
Agentic Performance (𝜏-Voice) by Domain
Note: Following models based on 1 trial: GPT-Realtime-2 (Minimal), OpenAI; Following models based on 2 trials: GPT-Realtime-2.1 High, OpenAI, OpenAI
Conversational Dynamics
Conversational Dynamics (Full Duplex Bench subset)
Conversational Dynamics (Full Duplex Bench subset) - Category Breakdown
API Benchmarks
Artificial Analysis Speech to Speech Index vs. Cost per Hour of Input Audio
Cost per Hour of Input Audio (Big Bench Audio subset)
Time to First Audio
Summary of Key Metrics & Further Information
Qwen Audio 3.0 Realtime Plus, Alibaba Cloud | 84.1% | 99% | 98.4% | 54.6% | 4.02 | 4.42 | 0.03 | 0.18 | |
Qwen3.5 Omni Plus Realtime | - | 99% | - | - | 2.64 | 0.00 | 0.16 | 0.82 | |
Qwen Audio 3.0 Realtime Flash, Alibaba Cloud | 76.3% | 96% | 96.9% | 35.9% | 4.16 | 4.77 | - | - | |
Qwen3.5 Omni Flash Realtime | - | 59% | - | - | 0.79 | 0.00 | 0.16 | 0.82 | |
Qwen3 Omni Flash | - | 59% | 72.7% | - | 4.82 | 1.77 | 0.61 | 1.36 | |
Qwen3 Omni Realtime | - | 57% | - | - | 0.88 | 2.26 | 0.65 | 0.99 | |
Nova 2.0 Sonic (Mar 2026) | - | 88% | - | - | 1.14 | - | 0.27 | 1.08 | |
FLM-Audio | - | 16% | 62.0% | - | - | - | - | - | |
Deepslate Opal | 62.8% | 85% | 85.7% | 17.5% | 0.44 | 6.48 | - | - | |
Gemini 3.1 Flash - High | 69.5% | 97% | 74.3% | 37.7% | 2.99 | 1.75 | 0.35 | 1.38 | |
Gemini 2.5 Flash Native Audio Dialog Thinking | - | 91% | - | - | 3.87 | - | 0.35 | 1.38 | |
Gemini 3.1 Flash - Minimal | 56.6% | 71% | 72.3% | 26.2% | 0.96 | 1.50 | 0.35 | 1.38 | |
Gemini 2.5 Flash Native Audio Dialog | - | 69% | - | - | 0.63 | 1.42 | 0.35 | 1.38 | |
Gemini 2.5 Flash Native Audio Preview (Dec 2025) | - | - | 44.0% | 22.8% | - | - | - | - | |
Gemini 2.5 Flash Native Audio Preview (Sep 2025) | - | - | 30.3% | - | - | - | - | - | |
Moshi | - | 4% | 61.0% | - | - | - | - | - | |
Nemotron Voicechat | - | 27% | 52.9% | - | - | - | - | - | |
PersonaPlex | - | 19% | 91.0% | - | - | - | - | - | |
GPT-Realtime-2 (High) | 77.2% | 97% | 95.3% | 39.8% | 1.14 | 4.14 | 1.15 | 4.61 | |
GPT-Realtime-2.1 High, OpenAI | 79.1% | 96% | 95.7% | 45.7% | 1.21 | 10.75 | - | - | |
GPT-Realtime-2 (Medium) | 75.3% | 93% | 95.2% | 37.4% | 1.22 | 3.97 | 1.15 | 4.61 | |
GPT-Realtime-2.1 Minimal, OpenAI | 72.5% | 87% | 92.7% | 38.0% | 0.97 | 11.31 | - | - | |
GPT Realtime | 69.2% | 83% | 93.9% | 30.4% | 0.98 | 11.08 | 1.15 | 4.61 | |
GPT-Realtime-1.5 | 72.0% | 81% | 95.7% | 38.8% | 0.81 | 11.44 | 1.15 | 4.61 | |
GPT-Realtime-2.1 Mini High, OpenAI | 65.3% | 75% | 91.7% | 29.4% | 4.28 | 3.45 | - | - | |
GPT-Realtime-2 (Minimal) | 66.2% | 72% | 96.1% | 30.8% | 1.12 | 3.07 | 1.15 | 4.61 | |
GPT-4o mini Realtime (Dec 2024) | - | 69% | - | - | 1.27 | 5.75 | 0.36 | 1.44 | |
GPT Realtime Mini (Oct 2025) | 58.1% | 64% | 95.7% | 15.1% | 0.81 | 3.04 | 0.36 | 1.44 | |
GPT-Realtime-2.1 Mini Minimal, OpenAI | 59.0% | 63% | 91.8% | 22.5% | 0.85 | 4.60 | - | - | |
GPT-4o audio chatcompletions | - | 54% | - | - | 3.38 | 0.00 | 3.60 | 14.40 | |
GPT-4o Realtime (Dec 2024) | - | - | 89.8% | 27.9% | 1.49 | 2.04 | 1.44 | 5.76 | |
Grok Voice Think Fast 2.0 High | 82.9% | 97% | 95.1% | 56.5% | 0.70 | 4.80 | - | - | |
Grok Voice Think Fast 1.0 | 75.7% | 97% | 77.8% | 52.1% | 1.25 | 3.00 | - | - | |
Grok Voice Fast 1.0 | 64.1% | 93% | 71.6% | 27.4% | 0.78 | 3.00 | 3.00 | 3.00 | |
Step-Audio R1.1 (Realtime) | - | 98% | - | - | 1.53 | 0.00 | 0.06 | 1.69 | |
Freeze-Omni | - | 33% | 58.7% | - | - | - | - | - |
Frequently Asked Questions
Qwen Audio 3.0 Realtime Plus, Alibaba Cloud leads with a speech reasoning score of 99.2% on the Big Bench Audio dataset across 33 models evaluated. Qwen3.5 Omni Plus Realtime follows at 98.7% and Step-Audio R1.1 (Realtime) at 97.6%.
Qwen Audio 3.0 Realtime Plus, Alibaba Cloud leads with a conversational dynamics score of 98.4% on the Full Duplex Bench dataset across 27 models evaluated. Qwen Audio 3.0 Realtime Flash, Alibaba Cloud follows at 96.9% and GPT-Realtime-2 (Minimal) at 96.1%.
The top speech to speech models by reasoning quality are: 1. Qwen Audio 3.0 Realtime Plus, Alibaba Cloud (99.2%), 2. Qwen3.5 Omni Plus Realtime (98.7%), 3. Step-Audio R1.1 (Realtime) (97.6%), 4. Grok Voice Think Fast 2.0 High (97.2%), 5. Grok Voice Think Fast 1.0 (97.1%). Scores are based on the Big Bench Audio benchmark.
The top speech to speech models by conversational dynamics are: 1. Qwen Audio 3.0 Realtime Plus, Alibaba Cloud (98.4%), 2. Qwen Audio 3.0 Realtime Flash, Alibaba Cloud (96.9%), 3. GPT-Realtime-2 (Minimal) (96.1%), 4. GPT-Realtime-2.1 High, OpenAI (95.7%), 5. GPT-Realtime-1.5 (95.7%). Scores are based on the Full Duplex Bench benchmark.
Deepslate Opal has the lowest Time to First Audio at 0.44s, followed by Gemini 2.5 Flash Native Audio Dialog (0.63s) and Grok Voice Think Fast 2.0 High (0.70s). Lower values mean faster initial response.
Qwen Audio 3.0 Realtime Plus, Alibaba Cloud is the most affordable at $0.0331/hour for input audio, followed by Step-Audio R1.1 (Realtime) ($0.064/hour) and Qwen3.5 Omni Flash Realtime ($0.162/hour).
Full Duplex Bench evaluates how well speech to speech models handle real conversational behaviors: knowing when to speak, when to stay silent during pauses, how to respond to interruptions, and how to recognize backchannels like "yeah" or "mm-hmm". Artificial Analysis implements a subset of Full Duplex Bench v1 and v1.5.
Speech reasoning (Big Bench Audio) measures whether a model can understand and correctly answer reasoning questions delivered as audio. Conversational dynamics (Full Duplex Bench) measures whether a model can handle natural conversation flow — turn-taking, pauses, and interruptions. A model may excel at one but not the other.
Real-time conversation requires both strong conversational dynamics and low latency. By conversational dynamics score, Qwen Audio 3.0 Realtime Plus, Alibaba Cloud (98.4%), Qwen Audio 3.0 Realtime Flash, Alibaba Cloud (96.9%), and GPT-Realtime-2 (Minimal) (96.1%) lead the field. By Time to First Audio, Deepslate Opal (0.44s), Gemini 2.5 Flash Native Audio Dialog (0.63s), and Grok Voice Think Fast 2.0 High (0.70s) are the fastest. The best choice depends on whether natural conversation flow or response speed is more critical for your use case.
Benchmarks are updated regularly as new models and providers are added. Performance metrics are continuously monitored to reflect current provider capabilities and pricing changes.