Ranking de modelos e provedores de IA de fala para texto

Compare a taxa de erro de palavras, a velocidade e os preços entre modelos e provedores de fala para texto.

Para mais detalhes, consulte nossa página de metodologia.

Destaques

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Seconds to final transcript after speech end · weighted average of samples in AA-WER Streaming · Lower is better
USD per 1000 minutes of audio · Lower is better

Índice AA-WER Streaming vs. tempo até a transcrição final

Índice AA-WER Streaming vs. tempo até a transcrição final

% de palavras transcritas incorretamente na transcrição final após o fim da fala detectado vs. segundos até a transcrição final após o fim da fala
Most attractive quadrant
Pareto line

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Índice AA-WER Streaming - Transcrição final

% de palavras transcritas incorretamente na transcrição final após o fim da fala detectado

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

AA-WER Streaming - Transcrição final: conjunto de dados AA-AgentTalk

% de palavras transcritas incorretamente na transcrição final após o fim da fala detectado no conjunto de dados AA-AgentTalk; menor é melhor

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Índice AA-WER Streaming (primeira parcial) vs. tempo até a primeira transcrição parcial após o fim da fala

Índice AA-WER Streaming (primeira parcial) vs. tempo até a primeira transcrição parcial após o fim da fala

% de palavras transcritas incorretamente na primeira transcrição parcial após o fim da fala detectado vs. segundos até a primeira transcrição parcial após o fim da fala
Most attractive quadrant
Pareto line

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

Índice AA-WER Streaming na primeira transcrição parcial após o fim da fala

% de palavras transcritas incorretamente na primeira transcrição parcial após o fim da fala detectado

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER Streaming na primeira transcrição parcial após o fim da fala: conjunto de dados AA-AgentTalk

% de palavras transcritas incorretamente na primeira transcrição parcial após o fim da fala detectado no conjunto de dados AA-AgentTalk; menor é melhor

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER Streaming - Transcrição final comparada com a primeira transcrição parcial após o fim da fala

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Latência

Tempo até a transcrição final

Segundos até a transcrição final após o fim da fala

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Tempo até a primeira transcrição parcial após o fim da fala

Segundos até a primeira transcrição parcial após o fim da fala

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

Preço

Preço da transcrição

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.