Zoom Scribe
Zoom Scribe

Zoom Scribe:API 服务商基准测试与分析

开发商:Zoom
许可证:专有
访问
分析 Zoom Scribe 的 API 服务商在 Artificial Analysis 词错误率指数、速度与价格等性能指标上的表现。

亮点

AA-WER v2 · % of words transcribed incorrectly · Lower is better
Input audio seconds transcribed per second · Higher is better
USD per 1000 minutes of audio · Lower is better

Artificial Analysis 词错误率(AA-WER)指数(按 API)

Artificial Analysis 词错误率(AA-WER)指数(按 API)

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%)
注意: 对于 Earnings22,如果模型因时间限制而无法可靠处理完整长度的音频,我们会将音频切分为约 9 分钟的片段(相关模型:Nova 2 Pro, AmazonGPT-4o Transcribe, OpenAIGPT-4o Mini Transcribe, OpenAI)。对于时间限制更短的模型,我们会将音频切分为约 30 秒的片段(相关模型:Canary Qwen 2.5B, NVIDIA)。

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

API 基准测试

Artificial Analysis 词错误率指数 vs. 价格

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%) · USD per 1000 minutes of audio
Most attractive quadrant

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

速度因子

速度因子

Input audio seconds transcribed per second · Higher is better

Audio file seconds transcribed per second of processing time. Higher factor indicates faster transcription speed. Reported Speed Factor values are medians across benchmark trials from the last 7 days; over-time chart points are daily medians. Artificial Analysis measurements are based on an audio duration of 10 minutes. Speed Factor may vary for other durations, particularly very short durations under 1 minute.

价格

转写价格

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

关键指标摘要与更多信息

服务商
更多详情
Qwen3.5 Omni Flash
Qwen3.5 Omni Flash 标志Alibaba Cloud
13.5%
78.5
0.00
Qwen3.5 Omni Plus
Qwen3.5 Omni Plus 标志Alibaba Cloud
3.5%
95.2
0.00
Nova 2 Pro
Nova 2 Pro 标志Amazon Bedrock
4.9%
23.0
3.10
Amazon Transcribe
Amazon Transcribe 标志Amazon Bedrock
4.1%
16.9
6.00
Universal-3 Pro
Universal-3 Pro 标志AssemblyAI
3.1%
111.7
3.50
Universal, AssemblyAI
Universal, AssemblyAI 标志AssemblyAI
3.8%
123.1
2.50
MAI-Transcribe-1.5
MAI-Transcribe-1.5 标志Microsoft Azure
2.4%
183.3
6.00
MAI-Transcribe-1
MAI-Transcribe-1 标志Microsoft Azure
2.6%
67.1
6.00
transcribe-03-2026
transcribe-03-2026 标志Cohere
4.6%
118.7
0.00
Nova-3
Nova-3 标志Deepgram
5.2%
541.1
4.30
Scribe v2
Scribe v2 标志ElevenLabs
2.2%
57.0
3.67
Solaria-1, Gladia
Solaria-1, Gladia 标志Gladia
4.1%
80.0
10.17
Solaria-3, Gladia
Solaria-3, Gladia 标志Gladia
3.2%
62.8
10.16
Gemini 3.1 Pro Preview (High)
Gemini 3.1 Pro Preview (High) 标志Google
2.8%
7.0
18.15
Gemini 3.1 Pro Preview (Low)
Gemini 3.1 Pro Preview (Low) 标志Google
3.6%
7.4
7.72
Gemini 3 Flash (High)
Gemini 3 Flash (High) 标志Google
2.9%
16.6
13.70
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite 标志Google
5.2%
69.0
6.56
Gemini 2.5 Flash
Gemini 2.5 Flash 标志Google
5.1%
74.5
6.66
Gemini 2.5 Pro
Gemini 2.5 Pro 标志Google
2.9%
13.3
11.39
Gemini 3.1 Flash-Lite Preview (Minimal)
Gemini 3.1 Flash-Lite Preview (Minimal) 标志Google
3.4%
82.3
5.83
Gradium Speech-to-Text
Gradium Speech-to-Text 标志Gradium
6.8%
2.3
13.00
Grok Speech to Text, SpaceXAI
Grok Speech to Text, SpaceXAI 标志SpaceXAI
4.0%
231.1
1.67
Voxtral Mini Transcribe 2
Voxtral Mini Transcribe 2 标志Mistral
3.6%
85.1
3.00
Voxtral Small
Voxtral Small 标志Mistral
2.8%
66.2
4.00
Voxtral Mini
Voxtral Mini 标志DeepInfra
3.8%
78.1
1.00
Modulate STT Batch English VFast
Modulate STT Batch English VFast 标志Modulate
4.2%
64.9
0.42
Parakeet TDT 0.6B V3, Togetherai
Parakeet TDT 0.6B V3, Togetherai 标志Together AI
4.5%
292.8
1.50
Canary Qwen 2.5B, NVIDIA
Canary Qwen 2.5B, NVIDIA 标志Replicate
4.3%
5.7
0.74
Parakeet TDT 0.6B V2, NVIDIA
Parakeet TDT 0.6B V2, NVIDIA 标志NVIDIA
6.4%
98.9
0.00
Parakeet RNNT 1.1B
Parakeet RNNT 1.1B 标志Replicate
5.4%
6.3
1.91
GPT Transcribe, OpenAI
GPT Transcribe, OpenAI 标志OpenAI
3.3%
38.9
4.50
GPT-4o Transcribe
GPT-4o Transcribe 标志OpenAI
4.0%
38.1
6.00
GPT-4o Mini Transcribe
GPT-4o Mini Transcribe 标志OpenAI
4.5%
43.5
3.00
Smallest AI Pulse Pro
Smallest AI Pulse Pro 标志Smallest.ai
2.4%
274.6
4.00
Resonant-1
Resonant-1 标志Reson8
3.4%
306.8
3.60
Rev AI
Rev AI 标志Rev AI
5.9%
13.4
3.33
Smallest AI Pulse
Smallest AI Pulse 标志Smallest.ai
4.4%
278.7
5.00
Soniox v5 Async
Soniox v5 Async 标志Soniox
3.8%
19.0
1.66
Soniox V4
Soniox V4 标志Soniox
3.9%
19.4
1.66
Speechmatics Melia
Speechmatics Melia 标志Speechmatics
4.9%
210.2
4.00
Speechmatics Standard
Speechmatics Standard 标志Speechmatics
5.1%
104.5
7.50
Speechmatics Enhanced
Speechmatics Enhanced 标志Speechmatics
4.0%
71.8
12.50
Whisper Large v3 Turbo
Whisper Large v3 Turbo 标志Groq
4.6%
117.1
0.67
Wizper Large v3
Wizper Large v3 标志fal.ai
4.7%
154.6
0.50
Incredibly Fast Whisper
Incredibly Fast Whisper 标志Replicate
5.7%
55.4
1.49
Whisper Large v3
Whisper Large v3 标志Replicate
10.1%
2.7
4.23
Whisper Large v3
Whisper Large v3 标志fal.ai
4.1%
89.2
1.15
Whisper Large v3
Whisper Large v3 标志Together AI
4.5%
311.8
1.50
Whisper Large v2
Whisper Large v2 标志OpenAI
4.1%
29.6
6.00