Zoom Scribe
Zoom Scribe

Zoom Scribe: APIプロバイダーのベンチマークと分析

開発元:Zoom
ライセンス:独自
訪問
Artificial Analysis単語誤り率インデックス、速度、料金を含むパフォーマンス指標でのZoom Scribe APIプロバイダーの分析。

注目情報

AA-WER v2 · % of words transcribed incorrectly · Lower is better
Input audio seconds transcribed per second · Higher is better
USD per 1000 minutes of audio · Lower is better

API別のArtificial Analysis単語誤り率(AA-WER)インデックス

API別のArtificial Analysis単語誤り率(AA-WER)インデックス

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%)
注: Earnings22で、時間制限によりモデルがフル尺の音声を安定して処理できない場合、約9分ごとに分割します(対象: Nova 2 Pro, Amazon GPT-4o Transcribe, OpenAI GPT-4o Mini Transcribe, OpenAI)。さらに短い時間制限があるモデルでは、約30秒ごとに分割します(対象: Canary Qwen 2.5B, NVIDIA)。

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

APIベンチマーク

Artificial Analysis単語誤り率インデックス vs. 料金

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%) · USD per 1000 minutes of audio
Most attractive quadrant

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

速度係数

速度係数

Input audio seconds transcribed per second · Higher is better

Audio file seconds transcribed per second of processing time. Higher factor indicates faster transcription speed. Reported Speed Factor values are medians across benchmark trials from the last 7 days; over-time chart points are daily medians. Artificial Analysis measurements are based on an audio duration of 10 minutes. Speed Factor may vary for other durations, particularly very short durations under 1 minute.

料金

文字起こし料金

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

主要指標の概要と詳細情報

プロバイダー
詳細情報
Qwen3.5 Omni Flash
Qwen3.5 Omni FlashのロゴAlibaba Cloud
13.5%
78.5
0.00
Qwen3.5 Omni Plus
Qwen3.5 Omni PlusのロゴAlibaba Cloud
3.5%
95.2
0.00
Nova 2 Pro
Nova 2 ProのロゴAmazon Bedrock
4.9%
23.0
3.10
Amazon Transcribe
Amazon TranscribeのロゴAmazon Bedrock
4.1%
16.9
6.00
Universal-3 Pro
Universal-3 ProのロゴAssemblyAI
3.1%
111.7
3.50
Universal, AssemblyAI
Universal, AssemblyAIのロゴAssemblyAI
3.8%
123.1
2.50
MAI-Transcribe-1.5
MAI-Transcribe-1.5のロゴMicrosoft Azure
2.4%
183.3
6.00
MAI-Transcribe-1
MAI-Transcribe-1のロゴMicrosoft Azure
2.6%
67.1
6.00
transcribe-03-2026
transcribe-03-2026のロゴCohere
4.6%
118.7
0.00
Nova-3
Nova-3のロゴDeepgram
5.2%
541.1
4.30
Scribe v2
Scribe v2のロゴElevenLabs
2.2%
57.0
3.67
Solaria-1, Gladia
Solaria-1, GladiaのロゴGladia
4.1%
80.0
10.17
Solaria-3, Gladia
Solaria-3, GladiaのロゴGladia
3.2%
62.8
10.16
Gemini 3.1 Pro Preview (High)
Gemini 3.1 Pro Preview (High)のロゴGoogle
2.8%
7.0
18.15
Gemini 3.1 Pro Preview (Low)
Gemini 3.1 Pro Preview (Low)のロゴGoogle
3.6%
7.4
7.72
Gemini 3 Flash (High)
Gemini 3 Flash (High)のロゴGoogle
2.9%
16.6
13.70
Gemini 2.5 Flash Lite
Gemini 2.5 Flash LiteのロゴGoogle
5.2%
69.0
6.56
Gemini 2.5 Flash
Gemini 2.5 FlashのロゴGoogle
5.1%
74.5
6.66
Gemini 2.5 Pro
Gemini 2.5 ProのロゴGoogle
2.9%
13.3
11.39
Gemini 3.1 Flash-Lite Preview (Minimal)
Gemini 3.1 Flash-Lite Preview (Minimal)のロゴGoogle
3.4%
82.3
5.83
Gradium Speech-to-Text
Gradium Speech-to-TextのロゴGradium
6.8%
2.3
13.00
Grok Speech to Text, SpaceXAI
Grok Speech to Text, SpaceXAIのロゴSpaceXAI
4.0%
231.1
1.67
Voxtral Mini Transcribe 2
Voxtral Mini Transcribe 2のロゴMistral
3.6%
85.1
3.00
Voxtral Small
Voxtral SmallのロゴMistral
2.8%
66.2
4.00
Voxtral Mini
Voxtral MiniのロゴDeepInfra
3.8%
78.1
1.00
Modulate STT Batch English VFast
Modulate STT Batch English VFastのロゴModulate
4.2%
64.9
0.42
Parakeet TDT 0.6B V3, Togetherai
Parakeet TDT 0.6B V3, TogetheraiのロゴTogether AI
4.5%
292.8
1.50
Canary Qwen 2.5B, NVIDIA
Canary Qwen 2.5B, NVIDIAのロゴReplicate
4.3%
5.7
0.74
Parakeet TDT 0.6B V2, NVIDIA
Parakeet TDT 0.6B V2, NVIDIAのロゴNVIDIA
6.4%
98.9
0.00
Parakeet RNNT 1.1B
Parakeet RNNT 1.1BのロゴReplicate
5.4%
6.3
1.91
GPT Transcribe, OpenAI
GPT Transcribe, OpenAIのロゴOpenAI
3.3%
38.9
4.50
GPT-4o Transcribe
GPT-4o TranscribeのロゴOpenAI
4.0%
38.1
6.00
GPT-4o Mini Transcribe
GPT-4o Mini TranscribeのロゴOpenAI
4.5%
43.5
3.00
Smallest AI Pulse Pro
Smallest AI Pulse ProのロゴSmallest.ai
2.4%
274.6
4.00
Resonant-1
Resonant-1のロゴReson8
3.4%
306.8
3.60
Rev AI
Rev AIのロゴRev AI
5.9%
13.4
3.33
Smallest AI Pulse
Smallest AI PulseのロゴSmallest.ai
4.4%
278.7
5.00
Soniox v5 Async
Soniox v5 AsyncのロゴSoniox
3.8%
19.0
1.66
Soniox V4
Soniox V4のロゴSoniox
3.9%
19.4
1.66
Speechmatics Melia
Speechmatics MeliaのロゴSpeechmatics
4.9%
210.2
4.00
Speechmatics Standard
Speechmatics StandardのロゴSpeechmatics
5.1%
104.5
7.50
Speechmatics Enhanced
Speechmatics EnhancedのロゴSpeechmatics
4.0%
71.8
12.50
Whisper Large v3 Turbo
Whisper Large v3 TurboのロゴGroq
4.6%
117.1
0.67
Wizper Large v3
Wizper Large v3のロゴfal.ai
4.7%
154.6
0.50
Incredibly Fast Whisper
Incredibly Fast WhisperのロゴReplicate
5.7%
55.4
1.49
Whisper Large v3
Whisper Large v3のロゴReplicate
10.1%
2.7
4.23
Whisper Large v3
Whisper Large v3のロゴfal.ai
4.1%
89.2
1.15
Whisper Large v3
Whisper Large v3のロゴTogether AI
4.5%
311.8
1.50
Whisper Large v2
Whisper Large v2のロゴOpenAI
4.1%
29.6
6.00