语音转文本 AI 模型与服务商排行榜

比较不同语音转文本模型和服务商的词错误率、速度与价格。

有关更多详细信息,请参阅我们的方法论页面

亮点

AA-WER v2 · % of words transcribed incorrectly · Lower is better
Input audio seconds transcribed per second · Higher is better
USD per 1000 minutes of audio · Lower is better

Artificial Analysis 词错误率指数(非流式)

Artificial Analysis 词错误率指数(非流式)

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%)
注意: 对于 Earnings22,如果模型因时间限制而无法可靠处理完整长度的音频,我们会将音频切分为约 9 分钟的片段(相关模型:GPT-4o Mini Transcribe, OpenAINova 2 Pro, AmazonGPT-4o Transcribe, OpenAI)。对于时间限制更短的模型,我们会将音频切分为约 30 秒的片段(相关模型:Inworld STT 1Canary Qwen 2.5B, NVIDIAQwen3 ASR Flash, Alibaba)。

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

AA-WER(非流式)按数据集

AA-WER(非流式):AA-AgentTalk 数据集

% of words transcribed incorrectly on the AA-AgentTalk dataset · Lower is better

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

清洗数据集对比

VoxPopuli:公开数据的清洗子集 vs. 原始子集

% WER (word error rate) · Lower is better
排序依据
注意:清洗版本会从参考文本中移除转写错误,从而为模型评估提供更准确的基准真值。

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

API 基准测试

Artificial Analysis 词错误率指数(非流式)vs. 价格

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%) · USD per 1000 minutes of audio
Most attractive quadrant

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

速度因子

Input audio seconds transcribed per second · Higher is better

Audio file seconds transcribed per second of processing time. Higher factor indicates faster transcription speed. Reported Speed Factor values are medians across benchmark trials from the last 7 days; over-time chart points are daily medians. Artificial Analysis measurements are based on an audio duration of 10 minutes. Speed Factor may vary for other durations, particularly very short durations under 1 minute.

转写价格

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

关键指标摘要与更多信息

服务商
更多详情
Qwen3.5 Omni Flash
Qwen3.5 Omni Flash 标志Alibaba Cloud
13.5%
77.7
0.00
Qwen3.5 Omni Plus
Qwen3.5 Omni Plus 标志Alibaba Cloud
3.5%
95.0
0.00
Nova 2 Pro
Nova 2 Pro 标志Amazon Bedrock
4.9%
22.9
3.10
Amazon Transcribe
Amazon Transcribe 标志Amazon Bedrock
4.1%
18.8
6.00
Universal-3 Pro
Universal-3 Pro 标志AssemblyAI
3.1%
98.7
3.50
Universal, AssemblyAI
Universal, AssemblyAI 标志AssemblyAI
3.8%
123.2
2.50
MAI-Transcribe-1.5
MAI-Transcribe-1.5 标志Microsoft Azure
2.4%
192.5
6.00
MAI-Transcribe-1
MAI-Transcribe-1 标志Microsoft Azure
2.6%
67.6
6.00
transcribe-03-2026
transcribe-03-2026 标志Cohere
4.6%
119.1
0.00
Nova-3
Nova-3 标志Deepgram
5.2%
545.1
4.30
Scribe v2
Scribe v2 标志ElevenLabs
2.2%
52.3
3.67
Solaria-1, Gladia
Solaria-1, Gladia 标志Gladia
4.1%
81.2
10.17
Solaria-3, Gladia
Solaria-3, Gladia 标志Gladia
3.2%
63.0
10.16
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe 标志Google
2.6%
79.6
5.00
Gemini 3.1 Pro Preview (High)
Gemini 3.1 Pro Preview (High) 标志Google
2.8%
7.3
18.15
Gemini 3.1 Pro Preview (Low)
Gemini 3.1 Pro Preview (Low) 标志Google
3.6%
6.5
7.72
Gemini 3 Flash (High)
Gemini 3 Flash (High) 标志Google
2.9%
18.3
13.70
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite 标志Google
5.2%
83.3
6.56
Gemini 2.5 Flash
Gemini 2.5 Flash 标志Google
5.1%
78.3
6.66
Gemini 2.5 Pro
Gemini 2.5 Pro 标志Google
2.9%
12.3
11.39
Gemini 3.1 Flash-Lite Preview (Minimal)
Gemini 3.1 Flash-Lite Preview (Minimal) 标志Google
3.4%
80.1
5.83
Gradium Speech-to-Text
Gradium Speech-to-Text 标志Gradium
6.8%
2.3
13.00
Grok Speech to Text, SpaceXAI
Grok Speech to Text, SpaceXAI 标志SpaceXAI
4.0%
226.4
1.67
Inworld STT 1
Inworld STT 1 标志Inworld
3.9%
206.1
2.50
Voxtral Mini Transcribe 2
Voxtral Mini Transcribe 2 标志Mistral
3.6%
83.1
3.00
Voxtral Small
Voxtral Small 标志Mistral
2.8%
66.3
4.00
Voxtral Mini
Voxtral Mini 标志DeepInfra
3.8%
79.3
1.00
Modulate STT Batch English VFast
Modulate STT Batch English VFast 标志Modulate
4.2%
61.9
0.42
Parakeet TDT 0.6B V3, Togetherai
Parakeet TDT 0.6B V3, Togetherai 标志Together AI
4.5%
274.9
1.50
Canary Qwen 2.5B, NVIDIA
Canary Qwen 2.5B, NVIDIA 标志Replicate
4.3%
6.7
0.74
Parakeet TDT 0.6B V2, NVIDIA
Parakeet TDT 0.6B V2, NVIDIA 标志NVIDIA
6.4%
100.1
0.00
Parakeet RNNT 1.1B
Parakeet RNNT 1.1B 标志Replicate
5.4%
6.3
1.91
GPT Transcribe, OpenAI
GPT Transcribe, OpenAI 标志OpenAI
3.3%
40.8
4.50
GPT-4o Transcribe
GPT-4o Transcribe 标志OpenAI
4.0%
37.1
6.00
GPT-4o Mini Transcribe
GPT-4o Mini Transcribe 标志OpenAI
4.5%
40.6
3.00
Smallest AI Pulse Pro
Smallest AI Pulse Pro 标志Smallest.ai
2.4%
272.4
4.00
Resonant-1
Resonant-1 标志Reson8
3.4%
332.2
3.60
Rev AI
Rev AI 标志Rev AI
5.9%
12.9
3.33
Smallest AI Pulse
Smallest AI Pulse 标志Smallest.ai
4.4%
278.3
5.00
Soniox v5 Async
Soniox v5 Async 标志Soniox
3.8%
36.8
1.66
Soniox V4
Soniox V4 标志Soniox
3.9%
39.7
1.66
Speechmatics Melia
Speechmatics Melia 标志Speechmatics
4.9%
206.8
4.00
Speechmatics Standard
Speechmatics Standard 标志Speechmatics
5.1%
107.1
7.50
Speechmatics Enhanced
Speechmatics Enhanced 标志Speechmatics
4.0%
73.3
12.50
StepAudio 2.5 ASR, StepFun
StepAudio 2.5 ASR, StepFun 标志StepFun
4.7%
81.8
0.37
Whisper Large v3 Turbo
Whisper Large v3 Turbo 标志Groq
4.6%
130.1
0.67
Wizper Large v3
Wizper Large v3 标志fal.ai
4.7%
287.6
0.50
Incredibly Fast Whisper
Incredibly Fast Whisper 标志Replicate
5.7%
55.0
1.49
Whisper Large v3
Whisper Large v3 标志Replicate
10.1%
2.6
4.23
Whisper Large v3
Whisper Large v3 标志fal.ai
4.1%
71.2
1.15
Whisper Large v3
Whisper Large v3 标志Together AI
4.5%
289.6
1.50
Whisper Large v2
Whisper Large v2 标志OpenAI
4.1%
29.0
6.00

常见问题

Fun-Realtime-ASR-preview 以最低的 AA-WER(Artificial Analysis 词错误率)1.7% 在 57 个已评估模型中领先。

按准确度(AA-WER)排名的顶级语音转文本模型为:1. Fun-Realtime-ASR-preview(1.7%)、2. Scribe v2, ElevenLabs(2.2%)、3. MAI-Transcribe-1.5(2.4%)、4. Smallest AI Pulse Pro(2.4%)和5. Gemini 3.5 Transcribe(2.6%)。AA-WER 越低,表示转写准确度越高。

Nova-3 最快,速度因子为实时的 545.1x,其次是 Resonant-1(332.2x)和 Whisper Large v3, together.ai(289.6x)。速度因子越高,转写越快。

StepAudio 2.5 ASR 最实惠,价格为每 1,000 分钟 $0.3667,其次是 Modulate STT Batch English VFast($0.417)和 Wizper (L, v3), fal.ai($0.50)。

Voxtral Small, Mistral 是最准确的开放权重模型,AA-WER 为 2.8%。在总共 57 个已评估模型中,有 12 个开放权重模型。

按准确度排名的顶级开放权重语音转文本模型为:1. Voxtral Small, Mistral(AA-WER 2.8%)、2. Inkling (256K), Thinking Machines(AA-WER 3.5%)和3. Voxtral Mini Transcribe 2, Mistral(AA-WER 3.6%)。

最佳模型取决于你的优先级。使用散点图可视化准确度(AA-WER)、速度与价格之间的权衡。对于需要高准确度的应用,优先选择 AA-WER 更低的模型;对于实时应用,关注速度因子;对于成本敏感的工作负载,比较价格图表。