Speech to Text AI Model & Provider Leaderboard
Compare word error rate, speed, and pricing across Speech to Text models and providers.
For further details, see our methodology page.
Highlights
Artificial Analysis Word Error Rate Index (Non-streaming)
Artificial Analysis Word Error Rate Index (Non-streaming)
AA-WER (Non-streaming) by Dataset
AA-WER (Non-streaming): AA-AgentTalk Dataset
Cleaned Dataset Comparison
VoxPopuli: Cleaned vs Original Subset of Publicly Available Data
API Benchmarks
Artificial Analysis Word Error Rate Index (Non-streaming) vs. Price
Speed Factor
Price of Transcription
Summary of Key Metrics & Further Information
Provider | Further Details | ||||
|---|---|---|---|---|---|
Qwen3.5 Omni Flash | 13.5% | 76.1 | 0.00 | ||
Qwen3.5 Omni Plus | 3.5% | 94.8 | 0.00 | ||
Nova 2 Pro | 4.9% | 22.8 | 3.10 | ||
Amazon Transcribe | 4.1% | 16.4 | 6.00 | ||
Universal-3 Pro | 3.1% | 91.3 | 3.50 | ||
Universal, AssemblyAI | 3.8% | 117.6 | 2.50 | ||
MAI-Transcribe-2 | 2.0% | 410.7 | 1.67 | ||
MAI-Transcribe-1.5 | 2.4% | 190.3 | 6.00 | ||
MAI-Transcribe-1 | 2.6% | 67.2 | 6.00 | ||
transcribe-03-2026 | 4.6% | 113.7 | 0.00 | ||
Nova-3 | 5.2% | 603.3 | 4.30 | ||
Nova-3, Telnyx | 4.8% | 360.9 | 7.40 | ||
Nova-2, Telnyx | 5.1% | 339.2 | 7.40 | ||
Scribe v2 | 2.2% | 53.8 | 3.67 | ||
Solaria-1, Gladia | 4.1% | 82.4 | 10.17 | ||
Solaria-3, Gladia | 3.2% | 63.6 | 10.16 | ||
Chirp 3, Google | 4.3% | 26.9 | 16.00 | ||
Chirp | 31.2% | 14.3 | 16.00 | ||
Gemini 3.5 Transcribe | 2.6% | 89.7 | 5.00 | ||
Gemini 3.1 Pro Preview (High) | 2.8% | 6.6 | 18.15 | ||
Gemini 3.1 Pro Preview (Low) | 3.6% | 6.0 | 7.72 | ||
Gemini 3 Flash (High) | 2.9% | 22.9 | 13.70 | ||
Gemini 2.5 Flash Lite | 5.2% | 83.6 | 6.56 | ||
Gemini 2.5 Flash | 5.1% | 83.7 | 6.66 | ||
Gemini 2.5 Pro | 2.9% | 11.3 | 11.39 | ||
Gemini 3.1 Flash-Lite Preview (Minimal) | 3.4% | 82.6 | 5.83 | ||
Gradium Speech-to-Text | 6.8% | 2.3 | 13.00 | ||
Grok Speech to Text, SpaceXAI | 4.0% | 219.9 | 1.67 | ||
Inworld STT 1 | 3.9% | 220.1 | 2.50 | ||
Voxtral Mini Transcribe 2 | 3.6% | 83.7 | 3.00 | ||
Voxtral Small | 2.8% | 65.7 | 4.00 | ||
Voxtral Mini | 3.8% | 77.1 | 1.00 | ||
Modulate STT Batch English VFast | 4.2% | 58.2 | 0.42 | ||
Canary Qwen 2.5B, NVIDIA | 4.3% | 6.9 | 0.74 | ||
Parakeet TDT 0.6B V2, NVIDIA | 6.4% | 100.3 | 0.00 | ||
Parakeet RNNT 1.1B | 5.4% | 6.3 | 1.91 | ||
GPT Transcribe, OpenAI | 3.3% | 40.0 | 4.50 | ||
GPT-4o Transcribe | 4.0% | 36.5 | 6.00 | ||
GPT-4o Mini Transcribe | 4.5% | 41.2 | 3.00 | ||
Smallest AI Pulse Pro | 2.4% | 274.4 | 4.00 | ||
Resonant-1 | 3.4% | 331.1 | 3.60 | ||
Rev AI | 5.9% | 11.9 | 3.33 | ||
Smallest AI Pulse | 4.4% | 284.5 | 5.00 | ||
Soniox v5 Async | 3.8% | 36.8 | 1.66 | ||
Soniox V4 | 3.9% | 44.7 | 1.66 | ||
Speechmatics Melia | 4.9% | 188.7 | 4.00 | ||
Speechmatics Standard | 5.1% | 97.6 | 7.50 | ||
Speechmatics Enhanced | 4.0% | 65.9 | 12.50 | ||
StepAudio 2.5 ASR, StepFun | 4.7% | 84.3 | 0.37 | ||
Whisper Large v3 Turbo | 4.6% | 153.5 | 0.67 | ||
Whisper Large v3 Turbo, Telnyx | 5.5% | 30.4 | 15.00 | ||
Wizper Large v3 | 4.7% | 181.4 | 0.50 | ||
Incredibly Fast Whisper | 5.7% | 54.4 | 1.49 | ||
Whisper Large v3 | 10.1% | 2.6 | 4.23 | ||
Whisper Large v3 | 4.1% | 103.8 | 1.15 | ||
Whisper Large v2 | 4.1% | 29.1 | 6.00 |
Frequently Asked Questions
Fun-Realtime-ASR-preview leads with the lowest AA-WER (Artificial Analysis Word Error Rate) of 1.7% across 58 models evaluated.
The top speech to text models by accuracy (AA-WER) are: 1. Fun-Realtime-ASR-preview (1.7%), 2. MAI-Transcribe-2 (2.0%), 3. Scribe v2, ElevenLabs (2.2%), 4. MAI-Transcribe-1.5 (2.4%), and 5. Smallest AI Pulse Pro (2.4%). Lower AA-WER indicates better transcription accuracy.
Nova-3 is the fastest with a speed factor of 603.3x real-time, followed by MAI-Transcribe-2 (410.7x) and Nova-3, Telnyx (360.9x). Higher speed factors mean faster transcription.
StepAudio 2.5 ASR is the most affordable at $0.3667 per 1,000 minutes, followed by Modulate STT Batch English VFast ($0.417) and Wizper (L, v3), fal.ai ($0.50).
Voxtral Small, Mistral is the most accurate open weights model with an AA-WER of 2.8%. There are 12 open weights models out of 58 total evaluated.
The top open weights speech to text models by accuracy are: 1. Voxtral Small, Mistral (AA-WER 2.8%), 2. Inkling (256K), Thinking Machines (AA-WER 3.5%), and 3. Voxtral Mini Transcribe 2, Mistral (AA-WER 3.6%).
The best model depends on your priorities. Use the scatter plots to visualize trade-offs between accuracy (AA-WER), speed, and price. For applications requiring high accuracy, prioritize models with lower AA-WER scores. For real-time applications, focus on speed factor. For cost-sensitive workloads, compare the price charts.