Speech to Text AI Model & Provider Leaderboard

Compare word error rate, speed, and pricing across Speech to Text models and providers.

For further details, see our methodology page.

Highlights

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Seconds to final transcript after speech end · weighted average of samples in AA-WER Streaming · Lower is better
USD per 1000 minutes of audio · Lower is better

AA-WER Streaming Index vs. Time to Final Transcription

AA-WER Streaming Index vs. Time to Final Transcription

% of words transcribed incorrectly at Final Transcription after detected End of Speech vs. Seconds to Final Transcription after End of Speech
Most attractive quadrant
Pareto line

AA-WER Streaming Index - Final Transcription

% of words transcribed incorrectly at Final Transcription after detected End of Speech

AA-WER Streaming - Final Transcription: AA-AgentTalk Dataset

% of words transcribed incorrectly at Final Transcription after detected End of Speech on AA-AgentTalk dataset; lower is better

AA-WER Streaming Index (First Partial) vs. Time to First Partial Transcription After Speech End

AA-WER Streaming Index (First Partial) vs. Time to First Partial Transcription After Speech End

% of words transcribed incorrectly at First Partial Transcription after detected End of Speech vs. Seconds to First Partial Transcription After Speech End
Most attractive quadrant
Pareto line

AA-WER Streaming Index at First Partial Transcription After Speech End

% of words transcribed incorrectly at First Partial Transcription after detected End of Speech

AA-WER Streaming at First Partial Transcription After Speech End: AA-AgentTalk Dataset

% of words transcribed incorrectly at First Partial Transcription after detected End of Speech on AA-AgentTalk dataset; lower is better

AA-WER Streaming - Final Transcription compared to First Partial Transcription After Speech End

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)

Latency

Time to Final Transcription

Seconds to Final Transcription after End of Speech

Time to First Partial Transcription After Speech End

Seconds to First Partial Transcription After Speech End

Price

Price of Transcription

USD per 1000 minutes of audio