Speech to Text AI Model & Provider Leaderboard

Compare word error rate, speed, and pricing across Speech to Text models and providers.

For further details, see our methodology page.

Highlights

AA-WER v2 · % of words transcribed incorrectly · Lower is better
Input audio seconds transcribed per second · Higher is better
USD per 1000 minutes of audio · Lower is better

Artificial Analysis Word Error Rate Index (Non-streaming)

Artificial Analysis Word Error Rate Index (Non-streaming)

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%)
Note: For Earnings22, if a model cannot reliably handle full-length audio due to time limits, we chunk to ~9 minutes (relevant to: GPT-4o Mini Transcribe, OpenAI, Nova 2 Pro, Amazon, GPT-4o Transcribe, OpenAI). For models with even shorter time limits, we chunk to ~30 seconds (relevant to: Inworld STT 1, Canary Qwen 2.5B, NVIDIA, Qwen3 ASR Flash, Alibaba).

AA-WER (Non-streaming) by Dataset

AA-WER (Non-streaming): AA-AgentTalk Dataset

% of words transcribed incorrectly on the AA-AgentTalk dataset · Lower is better

Cleaned Dataset Comparison

VoxPopuli: Cleaned vs Original Subset of Publicly Available Data

% WER (word error rate) · Lower is better
Sort by
Note: The cleaned versions remove transcription errors from the reference text, providing a more accurate ground truth for model evaluation.

API Benchmarks

Artificial Analysis Word Error Rate Index (Non-streaming) vs. Price

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%) · USD per 1000 minutes of audio
Most attractive quadrant

Speed Factor

Input audio seconds transcribed per second · Higher is better

Price of Transcription

USD per 1000 minutes of audio

Summary of Key Metrics & Further Information

Provider
Further Details
Qwen3.5 Omni Flash
Qwen3.5 Omni Flash logoAlibaba Cloud
13.5%
79.8
0.00
Qwen3.5 Omni Plus
Qwen3.5 Omni Plus logoAlibaba Cloud
3.5%
93.8
0.00
Nova 2 Pro
Nova 2 Pro logoAmazon Bedrock
4.9%
22.7
3.10
Amazon Transcribe
Amazon Transcribe logoAmazon Bedrock
4.1%
18.0
6.00
Universal-3 Pro
Universal-3 Pro logoAssemblyAI
3.1%
89.5
3.50
Universal, AssemblyAI
Universal, AssemblyAI logoAssemblyAI
3.8%
117.6
2.50
MAI-Transcribe-2
MAI-Transcribe-2 logoMicrosoft AI
2.0%
307.3
1.67
MAI-Transcribe-1.5
MAI-Transcribe-1.5 logoMicrosoft AI
2.4%
196.6
6.00
MAI-Transcribe-1
MAI-Transcribe-1 logoMicrosoft AI
2.6%
67.9
6.00
transcribe-03-2026
transcribe-03-2026 logoCohere
4.6%
123.2
0.00
Nova-3
Nova-3 logoDeepgram
5.2%
652.1
4.30
Nova-3, Telnyx
Nova-3, Telnyx logoTelnyx
4.8%
458.5
7.40
Nova-2, Telnyx
Nova-2, Telnyx logoTelnyx
5.1%
471.9
7.40
Scribe v2
Scribe v2 logoElevenLabs
2.2%
50.2
3.67
Solaria-1, Gladia
Solaria-1, Gladia logoGladia
4.1%
80.8
10.17
Solaria-3, Gladia
Solaria-3, Gladia logoGladia
3.2%
62.9
10.16
Chirp 3, Google
Chirp 3, Google logoGoogle
4.3%
28.7
16.00
Chirp
Chirp logoGoogle
31.2%
14.3
16.00
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe logoGoogle
2.6%
88.9
5.00
Gemini 3.1 Pro Preview (High)
Gemini 3.1 Pro Preview (High) logoGoogle
2.8%
6.7
18.15
Gemini 3.1 Pro Preview (Low)
Gemini 3.1 Pro Preview (Low) logoGoogle
3.6%
6.9
7.72
Gemini 3 Flash (High)
Gemini 3 Flash (High) logoGoogle
2.9%
22.6
13.70
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite logoGoogle
5.2%
82.6
6.56
Gemini 2.5 Flash
Gemini 2.5 Flash logoGoogle
5.1%
79.8
6.66
Gemini 2.5 Pro
Gemini 2.5 Pro logoGoogle
2.9%
12.4
11.39
Gemini 3.1 Flash-Lite Preview (Minimal)
Gemini 3.1 Flash-Lite Preview (Minimal) logoGoogle
3.4%
79.9
5.83
Gradium Speech-to-Text
Gradium Speech-to-Text logoGradium
6.8%
2.3
13.00
Grok Speech to Text, SpaceXAI
Grok Speech to Text, SpaceXAI logoSpaceXAI
4.0%
228.5
1.67
Inworld STT 1
Inworld STT 1 logoInworld
3.9%
207.5
2.50
Voxtral Mini Transcribe 2
Voxtral Mini Transcribe 2 logoMistral
3.6%
86.6
3.00
Voxtral Small
Voxtral Small logoMistral
2.8%
66.4
4.00
Voxtral Mini
Voxtral Mini logoDeepInfra
3.8%
79.8
1.00
Modulate STT Batch English VFast
Modulate STT Batch English VFast logoModulate
4.2%
59.8
0.42
Canary Qwen 2.5B, NVIDIA
Canary Qwen 2.5B, NVIDIA logoReplicate
4.3%
5.8
0.74
Parakeet TDT 0.6B V2, NVIDIA
Parakeet TDT 0.6B V2, NVIDIA logoNVIDIA
6.4%
98.9
0.00
Parakeet RNNT 1.1B
Parakeet RNNT 1.1B logoReplicate
5.4%
6.4
1.91
GPT Transcribe, OpenAI
GPT Transcribe, OpenAI logoOpenAI
3.3%
35.5
4.50
GPT-4o Transcribe
GPT-4o Transcribe logoOpenAI
4.0%
35.8
6.00
GPT-4o Mini Transcribe
GPT-4o Mini Transcribe logoOpenAI
4.5%
40.8
3.00
Smallest AI Pulse Pro
Smallest AI Pulse Pro logoSmallest.ai
2.4%
261.3
4.00
Resonant-1
Resonant-1 logoReson8
3.4%
329.0
3.60
Rev AI
Rev AI logoRev AI
5.9%
13.0
3.33
Smallest AI Pulse
Smallest AI Pulse logoSmallest.ai
4.4%
284.1
5.00
Soniox v5 Async
Soniox v5 Async logoSoniox
3.8%
36.7
1.66
Soniox V4
Soniox V4 logoSoniox
3.9%
42.1
1.66
Speechmatics Melia
Speechmatics Melia logoSpeechmatics
4.9%
178.4
4.00
Speechmatics Standard
Speechmatics Standard logoSpeechmatics
5.1%
96.7
7.50
Speechmatics Enhanced
Speechmatics Enhanced logoSpeechmatics
4.0%
75.4
12.50
StepAudio 2.5 ASR, StepFun
StepAudio 2.5 ASR, StepFun logoStepFun
4.7%
81.2
0.37
Whisper Large v3 Turbo
Whisper Large v3 Turbo logoGroq
4.6%
170.9
0.67
Whisper Large v3 Turbo, Telnyx
Whisper Large v3 Turbo, Telnyx logoTelnyx
5.5%
31.0
15.00
Wizper Large v3
Wizper Large v3 logofal.ai
4.7%
213.4
0.50
Incredibly Fast Whisper
Incredibly Fast Whisper logoReplicate
5.7%
56.2
1.49
Whisper Large v3
Whisper Large v3 logoReplicate
10.1%
2.8
4.23
Whisper Large v3
Whisper Large v3 logofal.ai
4.1%
55.0
1.15
Whisper Large v2
Whisper Large v2 logoOpenAI
4.1%
28.5
6.00

Frequently Asked Questions