음성 텍스트 변환 AI 모델 및 제공업체 리더보드

음성 텍스트 변환 모델과 제공업체의 단어 오류율, 속도, 가격을 비교합니다.

자세한 내용은 방법론 페이지를 참조하세요.

주요 내용

AA-WER v2 · % of words transcribed incorrectly · Lower is better
Input audio seconds transcribed per second · Higher is better
USD per 1000 minutes of audio · Lower is better

Artificial Analysis 단어 오류율 지수(비스트리밍)

Artificial Analysis 단어 오류율 지수(비스트리밍)

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%)
참고: Earnings22에서 시간 제한으로 인해 모델이 전체 길이의 오디오를 안정적으로 처리할 수 없는 경우 약 9분 단위로 분할합니다(해당 모델: GPT-4o Mini Transcribe, OpenAI Nova 2 Pro, Amazon GPT-4o Transcribe, OpenAI). 시간 제한이 더 짧은 모델은 약 30초 단위로 분할합니다(해당 모델: Inworld STT 1 Canary Qwen 2.5B, NVIDIA Qwen3 ASR Flash, Alibaba).

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

AA-WER(비스트리밍) 데이터 세트별

AA-WER(비스트리밍): AA-AgentTalk 데이터 세트

% of words transcribed incorrectly on the AA-AgentTalk dataset · Lower is better

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

정제 데이터 세트 비교

VoxPopuli: 공개 데이터의 정제 vs. 원본 하위 집합

% WER (word error rate) · Lower is better
정렬 기준
참고: 정제 버전은 참조 텍스트에서 전사 오류를 제거하여 모델 평가를 위한 더 정확한 정답을 제공합니다.

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

API 벤치마크

Artificial Analysis 단어 오류율 지수(비스트리밍) vs. 가격

% of words transcribed incorrectly · Lower is better · AA-WER v2 incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), Earnings22-Cleaned-AA (25%) · USD per 1000 minutes of audio
Most attractive quadrant

Measures transcription accuracy across 3 datasets to evaluate models in real-world speech with diverse accents, domain-specific language, and challenging channel & acoustic conditions.

AA-WER is calculated as an audio-duration-weighted average of WER across ~8 hours from three datasets: AA-AgentTalk (50%), VoxPopuli-Cleaned-AA (25%), and Earnings22-Cleaned-AA (25%). See methodology for more detail.

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

속도 계수

Input audio seconds transcribed per second · Higher is better

Audio file seconds transcribed per second of processing time. Higher factor indicates faster transcription speed. Reported Speed Factor values are medians across benchmark trials from the last 7 days; over-time chart points are daily medians. Artificial Analysis measurements are based on an audio duration of 10 minutes. Speed Factor may vary for other durations, particularly very short durations under 1 minute.

전사 가격

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.

주요 지표 요약 및 추가 정보

제공업체
추가 세부 정보
Qwen3.5 Omni Flash
Qwen3.5 Omni Flash 로고Alibaba Cloud
13.5%
77.9
0.00
Qwen3.5 Omni Plus
Qwen3.5 Omni Plus 로고Alibaba Cloud
3.5%
95.0
0.00
Nova 2 Pro
Nova 2 Pro 로고Amazon Bedrock
4.9%
22.9
3.10
Amazon Transcribe
Amazon Transcribe 로고Amazon Bedrock
4.1%
18.8
6.00
Universal-3 Pro
Universal-3 Pro 로고AssemblyAI
3.1%
98.7
3.50
Universal, AssemblyAI
Universal, AssemblyAI 로고AssemblyAI
3.8%
123.3
2.50
MAI-Transcribe-1.5
MAI-Transcribe-1.5 로고Microsoft Azure
2.4%
192.5
6.00
MAI-Transcribe-1
MAI-Transcribe-1 로고Microsoft Azure
2.6%
67.8
6.00
transcribe-03-2026
transcribe-03-2026 로고Cohere
4.6%
119.1
0.00
Nova-3
Nova-3 로고Deepgram
5.2%
501.4
4.30
Scribe v2
Scribe v2 로고ElevenLabs
2.2%
53.9
3.67
Solaria-1, Gladia
Solaria-1, Gladia 로고Gladia
4.1%
81.2
10.17
Solaria-3, Gladia
Solaria-3, Gladia 로고Gladia
3.2%
62.7
10.16
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe 로고Google
2.6%
82.5
5.00
Gemini 3.1 Pro Preview (High)
Gemini 3.1 Pro Preview (High) 로고Google
2.8%
7.4
18.15
Gemini 3.1 Pro Preview (Low)
Gemini 3.1 Pro Preview (Low) 로고Google
3.6%
6.5
7.72
Gemini 3 Flash (High)
Gemini 3 Flash (High) 로고Google
2.9%
18.3
13.70
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite 로고Google
5.2%
81.2
6.56
Gemini 2.5 Flash
Gemini 2.5 Flash 로고Google
5.1%
77.3
6.66
Gemini 2.5 Pro
Gemini 2.5 Pro 로고Google
2.9%
12.3
11.39
Gemini 3.1 Flash-Lite Preview (Minimal)
Gemini 3.1 Flash-Lite Preview (Minimal) 로고Google
3.4%
79.9
5.83
Gradium Speech-to-Text
Gradium Speech-to-Text 로고Gradium
6.8%
2.3
13.00
Grok Speech to Text, SpaceXAI
Grok Speech to Text, SpaceXAI 로고SpaceXAI
4.0%
225.0
1.67
Inworld STT 1
Inworld STT 1 로고Inworld
3.9%
206.9
2.50
Voxtral Mini Transcribe 2
Voxtral Mini Transcribe 2 로고Mistral
3.6%
82.6
3.00
Voxtral Small
Voxtral Small 로고Mistral
2.8%
65.7
4.00
Voxtral Mini
Voxtral Mini 로고DeepInfra
3.8%
79.1
1.00
Modulate STT Batch English VFast
Modulate STT Batch English VFast 로고Modulate
4.2%
61.9
0.42
Parakeet TDT 0.6B V3, Togetherai
Parakeet TDT 0.6B V3, Togetherai 로고Together AI
4.5%
273.1
1.50
Canary Qwen 2.5B, NVIDIA
Canary Qwen 2.5B, NVIDIA 로고Replicate
4.3%
7.9
0.74
Parakeet TDT 0.6B V2, NVIDIA
Parakeet TDT 0.6B V2, NVIDIA 로고NVIDIA
6.4%
99.9
0.00
Parakeet RNNT 1.1B
Parakeet RNNT 1.1B 로고Replicate
5.4%
6.3
1.91
GPT Transcribe, OpenAI
GPT Transcribe, OpenAI 로고OpenAI
3.3%
40.8
4.50
GPT-4o Transcribe
GPT-4o Transcribe 로고OpenAI
4.0%
37.1
6.00
GPT-4o Mini Transcribe
GPT-4o Mini Transcribe 로고OpenAI
4.5%
41.4
3.00
Smallest AI Pulse Pro
Smallest AI Pulse Pro 로고Smallest.ai
2.4%
272.9
4.00
Resonant-1
Resonant-1 로고Reson8
3.4%
331.6
3.60
Rev AI
Rev AI 로고Rev AI
5.9%
12.9
3.33
Smallest AI Pulse
Smallest AI Pulse 로고Smallest.ai
4.4%
274.9
5.00
Soniox v5 Async
Soniox v5 Async 로고Soniox
3.8%
35.1
1.66
Soniox V4
Soniox V4 로고Soniox
3.9%
39.7
1.66
Speechmatics Melia
Speechmatics Melia 로고Speechmatics
4.9%
204.1
4.00
Speechmatics Standard
Speechmatics Standard 로고Speechmatics
5.1%
107.1
7.50
Speechmatics Enhanced
Speechmatics Enhanced 로고Speechmatics
4.0%
70.6
12.50
StepAudio 2.5 ASR, StepFun
StepAudio 2.5 ASR, StepFun 로고StepFun
4.7%
81.5
0.37
Whisper Large v3 Turbo
Whisper Large v3 Turbo 로고Groq
4.6%
130.1
0.67
Wizper Large v3
Wizper Large v3 로고fal.ai
4.7%
287.6
0.50
Incredibly Fast Whisper
Incredibly Fast Whisper 로고Replicate
5.7%
55.1
1.49
Whisper Large v3
Whisper Large v3 로고Replicate
10.1%
2.6
4.23
Whisper Large v3
Whisper Large v3 로고fal.ai
4.1%
71.8
1.15
Whisper Large v3
Whisper Large v3 로고Together AI
4.5%
288.1
1.50
Whisper Large v2
Whisper Large v2 로고OpenAI
4.1%
28.8
6.00

자주 묻는 질문

Fun-Realtime-ASR-preview이(가) 평가된 57개 모델 중 가장 낮은 AA-WER(Artificial Analysis 단어 오류율) 1.7%로 선두입니다.

정확도(AA-WER) 기준 상위 Speech to Text 모델은 다음과 같습니다: 1. Fun-Realtime-ASR-preview (1.7%), 2. Scribe v2, ElevenLabs (2.2%), 3. MAI-Transcribe-1.5 (2.4%), 4. Smallest AI Pulse Pro (2.4%) 및 5. Gemini 3.5 Transcribe (2.6%). AA-WER이 낮을수록 전사 정확도가 더 높습니다.

Nova-3이(가) 실시간 대비 501.4x의 속도 계수로 가장 빠르며, 그다음이 Resonant-1(331.6x), Whisper Large v3, together.ai(288.1x)입니다. 속도 계수가 높을수록 전사가 더 빠릅니다.

StepAudio 2.5 ASR이(가) 1,000분당 $0.3667로 가장 저렴하며, 그다음이 Modulate STT Batch English VFast($0.417), Wizper (L, v3), fal.ai($0.50)입니다.

Voxtral Small, Mistral이(가) AA-WER 2.8%로 가장 정확한 오픈 웨이트 모델입니다. 평가된 전체 57개 중 오픈 웨이트 모델은 12개입니다.

정확도 기준 상위 오픈 웨이트 Speech to Text 모델은 다음과 같습니다: 1. Voxtral Small, Mistral (AA-WER 2.8%), 2. Inkling (256K), Thinking Machines (AA-WER 3.5%) 및 3. Voxtral Mini Transcribe 2, Mistral (AA-WER 3.6%).

최고의 모델은 우선순위에 따라 달라집니다. 산점도를 사용해 정확도(AA-WER), 속도, 가격 간의 절충을 시각화하세요. 높은 정확도가 필요한 애플리케이션에서는 AA-WER이 낮은 모델을 우선하고, 실시간 애플리케이션에서는 속도 계수에 집중하며, 비용에 민감한 워크로드에서는 가격 차트를 비교하세요.