Tabla de clasificación de modelos y proveedores de voz a texto

Compara la tasa de error de palabras, la velocidad y los precios entre modelos y proveedores de voz a texto.

Para más detalles, consulta nuestra página de metodología.

Destacados

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Seconds to final transcript after speech end · weighted average of samples in AA-WER Streaming · Lower is better
USD per 1000 minutes of audio · Lower is better

Índice AA-WER Streaming vs. tiempo hasta la transcripción final

Índice AA-WER Streaming vs. tiempo hasta la transcripción final

% de palabras transcritas incorrectamente en la transcripción final tras el final del habla detectado vs. segundos hasta la transcripción final tras el final del habla
Most attractive quadrant
Pareto line

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Índice AA-WER Streaming - Transcripción final

% de palabras transcritas incorrectamente en la transcripción final tras el final del habla detectado

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

AA-WER Streaming - Transcripción final: conjunto de datos AA-AgentTalk

% de palabras transcritas incorrectamente en la transcripción final tras el final del habla detectado en el conjunto de datos AA-AgentTalk; menor es mejor

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Índice AA-WER Streaming (primera parcial) vs. tiempo hasta la primera transcripción parcial tras el final del habla

Índice AA-WER Streaming (primera parcial) vs. tiempo hasta la primera transcripción parcial tras el final del habla

% de palabras transcritas incorrectamente en la primera transcripción parcial tras el final del habla detectado vs. segundos hasta la primera transcripción parcial tras el final del habla
Most attractive quadrant
Pareto line

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

Índice AA-WER Streaming en la primera transcripción parcial tras el final del habla

% de palabras transcritas incorrectamente en la primera transcripción parcial tras el final del habla detectado

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER Streaming en la primera transcripción parcial tras el final del habla: conjunto de datos AA-AgentTalk

% de palabras transcritas incorrectamente en la primera transcripción parcial tras el final del habla detectado en el conjunto de datos AA-AgentTalk; menor es mejor

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER Streaming - Transcripción final comparada con la primera transcripción parcial tras el final del habla

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Latencia

Tiempo hasta la transcripción final

Segundos hasta la transcripción final tras el final del habla

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Tiempo hasta la primera transcripción parcial tras el final del habla

Segundos hasta la primera transcripción parcial tras el final del habla

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

Precio

Precio de transcripción

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.