音声文字起こしAIモデル・プロバイダーランキング

音声文字起こしモデルとプロバイダーの単語誤り率、速度、料金を比較します。

詳しくは方法論ページをご覧ください。

注目情報

AA-WER v2 · % of words transcribed incorrectly · Lower is better
Input audio seconds transcribed per second · Higher is better
USD per 1000 minutes of audio · Lower is better

Artificial Analysis単語誤り率インデックス(非ストリーミング)

Artificial Analysis単語誤り率インデックス(非ストリーミング)

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%)
注: Earnings22で、時間制限によりモデルがフル尺の音声を安定して処理できない場合、約9分ごとに分割します(対象: GPT-4o Mini Transcribe, OpenAI Nova 2 Pro, Amazon GPT-4o Transcribe, OpenAI)。さらに短い時間制限があるモデルでは、約30秒ごとに分割します(対象: Inworld STT 1 Canary Qwen 2.5B, NVIDIA Qwen3 ASR Flash, Alibaba)。

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

AA-WER(非ストリーミング)データセット別

AA-WER(非ストリーミング): AA-AgentTalkデータセット

% of words transcribed incorrectly on the AA-AgentTalk dataset · Lower is better

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

クリーン済みデータセットの比較

VoxPopuli: 公開データのクリーン済みサブセット vs. オリジナル

% WER (word error rate) · Lower is better
並び替え
注: クリーン済み版は参照テキストから文字起こし誤りを除去し、モデル評価のためのより正確な正解データを提供します。

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

APIベンチマーク

Artificial Analysis単語誤り率インデックス(非ストリーミング)vs. 料金

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%) · USD per 1000 minutes of audio
Most attractive quadrant

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

速度係数

Input audio seconds transcribed per second · Higher is better

Audio file seconds transcribed per second of processing time. Higher factor indicates faster transcription speed. Reported Speed Factor values are medians across benchmark trials from the last 7 days; over-time chart points are daily medians. Artificial Analysis measurements are based on an audio duration of 10 minutes. Speed Factor may vary for other durations, particularly very short durations under 1 minute.

文字起こし料金

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

主要指標の概要と詳細情報

プロバイダー
詳細情報
Qwen3.5 Omni Flash
Qwen3.5 Omni FlashのロゴAlibaba Cloud
13.5%
77.9
0.00
Qwen3.5 Omni Plus
Qwen3.5 Omni PlusのロゴAlibaba Cloud
3.5%
95.0
0.00
Nova 2 Pro
Nova 2 ProのロゴAmazon Bedrock
4.9%
22.9
3.10
Amazon Transcribe
Amazon TranscribeのロゴAmazon Bedrock
4.1%
18.8
6.00
Universal-3 Pro
Universal-3 ProのロゴAssemblyAI
3.1%
98.7
3.50
Universal, AssemblyAI
Universal, AssemblyAIのロゴAssemblyAI
3.8%
123.3
2.50
MAI-Transcribe-1.5
MAI-Transcribe-1.5のロゴMicrosoft Azure
2.4%
192.5
6.00
MAI-Transcribe-1
MAI-Transcribe-1のロゴMicrosoft Azure
2.6%
67.8
6.00
transcribe-03-2026
transcribe-03-2026のロゴCohere
4.6%
119.1
0.00
Nova-3
Nova-3のロゴDeepgram
5.2%
501.4
4.30
Scribe v2
Scribe v2のロゴElevenLabs
2.2%
53.9
3.67
Solaria-1, Gladia
Solaria-1, GladiaのロゴGladia
4.1%
81.2
10.17
Solaria-3, Gladia
Solaria-3, GladiaのロゴGladia
3.2%
62.7
10.16
Gemini 3.5 Transcribe
Gemini 3.5 TranscribeのロゴGoogle
2.6%
82.5
5.00
Gemini 3.1 Pro Preview (High)
Gemini 3.1 Pro Preview (High)のロゴGoogle
2.8%
7.4
18.15
Gemini 3.1 Pro Preview (Low)
Gemini 3.1 Pro Preview (Low)のロゴGoogle
3.6%
6.5
7.72
Gemini 3 Flash (High)
Gemini 3 Flash (High)のロゴGoogle
2.9%
18.3
13.70
Gemini 2.5 Flash Lite
Gemini 2.5 Flash LiteのロゴGoogle
5.2%
81.2
6.56
Gemini 2.5 Flash
Gemini 2.5 FlashのロゴGoogle
5.1%
77.3
6.66
Gemini 2.5 Pro
Gemini 2.5 ProのロゴGoogle
2.9%
12.3
11.39
Gemini 3.1 Flash-Lite Preview (Minimal)
Gemini 3.1 Flash-Lite Preview (Minimal)のロゴGoogle
3.4%
79.9
5.83
Gradium Speech-to-Text
Gradium Speech-to-TextのロゴGradium
6.8%
2.3
13.00
Grok Speech to Text, SpaceXAI
Grok Speech to Text, SpaceXAIのロゴSpaceXAI
4.0%
225.0
1.67
Inworld STT 1
Inworld STT 1のロゴInworld
3.9%
206.9
2.50
Voxtral Mini Transcribe 2
Voxtral Mini Transcribe 2のロゴMistral
3.6%
82.6
3.00
Voxtral Small
Voxtral SmallのロゴMistral
2.8%
65.7
4.00
Voxtral Mini
Voxtral MiniのロゴDeepInfra
3.8%
79.1
1.00
Modulate STT Batch English VFast
Modulate STT Batch English VFastのロゴModulate
4.2%
61.9
0.42
Parakeet TDT 0.6B V3, Togetherai
Parakeet TDT 0.6B V3, TogetheraiのロゴTogether AI
4.5%
273.1
1.50
Canary Qwen 2.5B, NVIDIA
Canary Qwen 2.5B, NVIDIAのロゴReplicate
4.3%
7.9
0.74
Parakeet TDT 0.6B V2, NVIDIA
Parakeet TDT 0.6B V2, NVIDIAのロゴNVIDIA
6.4%
99.9
0.00
Parakeet RNNT 1.1B
Parakeet RNNT 1.1BのロゴReplicate
5.4%
6.3
1.91
GPT Transcribe, OpenAI
GPT Transcribe, OpenAIのロゴOpenAI
3.3%
40.8
4.50
GPT-4o Transcribe
GPT-4o TranscribeのロゴOpenAI
4.0%
37.1
6.00
GPT-4o Mini Transcribe
GPT-4o Mini TranscribeのロゴOpenAI
4.5%
41.4
3.00
Smallest AI Pulse Pro
Smallest AI Pulse ProのロゴSmallest.ai
2.4%
272.9
4.00
Resonant-1
Resonant-1のロゴReson8
3.4%
331.6
3.60
Rev AI
Rev AIのロゴRev AI
5.9%
12.9
3.33
Smallest AI Pulse
Smallest AI PulseのロゴSmallest.ai
4.4%
274.9
5.00
Soniox v5 Async
Soniox v5 AsyncのロゴSoniox
3.8%
35.1
1.66
Soniox V4
Soniox V4のロゴSoniox
3.9%
39.7
1.66
Speechmatics Melia
Speechmatics MeliaのロゴSpeechmatics
4.9%
204.1
4.00
Speechmatics Standard
Speechmatics StandardのロゴSpeechmatics
5.1%
107.1
7.50
Speechmatics Enhanced
Speechmatics EnhancedのロゴSpeechmatics
4.0%
70.6
12.50
StepAudio 2.5 ASR, StepFun
StepAudio 2.5 ASR, StepFunのロゴStepFun
4.7%
81.5
0.37
Whisper Large v3 Turbo
Whisper Large v3 TurboのロゴGroq
4.6%
130.1
0.67
Wizper Large v3
Wizper Large v3のロゴfal.ai
4.7%
287.6
0.50
Incredibly Fast Whisper
Incredibly Fast WhisperのロゴReplicate
5.7%
55.1
1.49
Whisper Large v3
Whisper Large v3のロゴReplicate
10.1%
2.6
4.23
Whisper Large v3
Whisper Large v3のロゴfal.ai
4.1%
71.8
1.15
Whisper Large v3
Whisper Large v3のロゴTogether AI
4.5%
288.1
1.50
Whisper Large v2
Whisper Large v2のロゴOpenAI
4.1%
28.8
6.00

よくある質問

Fun-Realtime-ASR-previewは、評価した57モデルのうち最も低いAA-WER(Artificial Analysis単語誤り率)1.7%で首位です。

精度(AA-WER)に基づく上位のSpeech to Textモデルは次のとおりです: 1. Fun-Realtime-ASR-preview(1.7%)、2. Scribe v2, ElevenLabs(2.2%)、3. MAI-Transcribe-1.5(2.4%)、4. Smallest AI Pulse Pro(2.4%)、5. Gemini 3.5 Transcribe(2.6%)。AA-WERが低いほど文字起こし精度が高いことを示します。

Nova-3が最速で、リアルタイムの501.4xの速度係数です。続いてResonant-1(331.6x)、Whisper Large v3, together.ai(288.1x)です。速度係数が高いほど文字起こしが高速です。

StepAudio 2.5 ASRが最も安価で、1,000分あたり$0.3667です。続いてModulate STT Batch English VFast($0.417)、Wizper (L, v3), fal.ai($0.50)です。

Voxtral Small, Mistralは最も精度の高いオープンウェイトモデルで、AA-WERは2.8%です。評価した合計57モデルのうち、オープンウェイトは12モデルです。

精度に基づく上位のオープンウェイトSpeech to Textモデルは次のとおりです: 1. Voxtral Small, Mistral(AA-WER 2.8%)、2. Inkling (256K), Thinking Machines(AA-WER 3.5%)、3. Voxtral Mini Transcribe 2, Mistral(AA-WER 3.6%)。

最適なモデルは優先事項によって異なります。散布図を使って精度(AA-WER)、速度、料金のトレードオフを可視化してください。高い精度が必要な用途ではAA-WERが低いモデルを優先し、リアルタイム用途では速度係数に注目し、コスト重視のワークロードでは料金チャートを比較してください。