Speech to Text AI Model & Provider Leaderboard
Compare word error rate, speed, and pricing across Speech to Text models and providers.
For further details, see our methodology page.
Highlights
% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Seconds to final transcript after speech end · weighted average of samples in AA-WER Streaming · Lower is better
AA-WER Streaming Index vs. Time to Final Transcription
AA-WER Streaming Index vs. Time to Final Transcription
% of words transcribed incorrectly at Final Transcription after detected End of Speech vs. Seconds to Final Transcription after End of Speech
Most attractive quadrant
Pareto line
AA-WER Streaming Index - Final Transcription
% of words transcribed incorrectly at Final Transcription after detected End of Speech
AA-WER Streaming - Final Transcription: AA-AgentTalk Dataset
% of words transcribed incorrectly at Final Transcription after detected End of Speech on AA-AgentTalk dataset; lower is better
AA-WER Streaming Index (First Partial) vs. Time to First Partial Transcription After Speech End
AA-WER Streaming Index (First Partial) vs. Time to First Partial Transcription After Speech End
% of words transcribed incorrectly at First Partial Transcription after detected End of Speech vs. Seconds to First Partial Transcription After Speech End
Most attractive quadrant
Pareto line
AA-WER Streaming Index at First Partial Transcription After Speech End
% of words transcribed incorrectly at First Partial Transcription after detected End of Speech
AA-WER Streaming at First Partial Transcription After Speech End: AA-AgentTalk Dataset
% of words transcribed incorrectly at First Partial Transcription after detected End of Speech on AA-AgentTalk dataset; lower is better
AA-WER Streaming - Final Transcription compared to First Partial Transcription After Speech End
% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Latency
Time to Final Transcription
Seconds to Final Transcription after End of Speech
Time to First Partial Transcription After Speech End
Seconds to First Partial Transcription After Speech End
Price
Price of Transcription
USD per 1000 minutes of audio