Bestenliste für Speech-to-Text-KI-Modelle und -Anbieter

Vergleichen Sie Wortfehlerrate, Geschwindigkeit und Preise von Speech-to-Text-Modellen und -Anbietern.

Weitere Einzelheiten finden Sie auf unserer Methodikseite.

Highlights

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Seconds to final transcript after speech end · weighted average of samples in AA-WER Streaming · Lower is better
USD per 1000 minutes of audio · Lower is better

AA-WER-Streaming-Index vs. Zeit bis zur finalen Transkription

AA-WER-Streaming-Index vs. Zeit bis zur finalen Transkription

% falsch transkribierter Wörter bei der finalen Transkription nach erkanntem Sprachende vs. Sekunden bis zur finalen Transkription nach Sprachende
Most attractive quadrant
Pareto line

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

AA-WER-Streaming-Index - Finale Transkription

% falsch transkribierter Wörter bei der finalen Transkription nach erkanntem Sprachende

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

AA-WER Streaming - Finale Transkription: AA-AgentTalk-Datensatz

% falsch transkribierter Wörter bei der finalen Transkription nach erkanntem Sprachende im AA-AgentTalk-Datensatz; niedriger ist besser

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

AA-WER-Streaming-Index (erste partielle) vs. Zeit bis zur ersten partiellen Transkription nach Sprachende

AA-WER-Streaming-Index (erste partielle) vs. Zeit bis zur ersten partiellen Transkription nach Sprachende

% falsch transkribierter Wörter bei der ersten partiellen Transkription nach erkanntem Sprachende vs. Sekunden bis zur ersten partiellen Transkription nach Sprachende
Most attractive quadrant
Pareto line

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER-Streaming-Index bei der ersten partiellen Transkription nach Sprachende

% falsch transkribierter Wörter bei der ersten partiellen Transkription nach erkanntem Sprachende

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER Streaming bei der ersten partiellen Transkription nach Sprachende: AA-AgentTalk-Datensatz

% falsch transkribierter Wörter bei der ersten partiellen Transkription nach erkanntem Sprachende im AA-AgentTalk-Datensatz; niedriger ist besser

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

AA-WER Streaming - Finale Transkription im Vergleich zur ersten partiellen Transkription nach Sprachende

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)

Measures transcription accuracy of models where audio is streamed in real-time, chunk by chunk, as opposed to batch transcription, where the full audio file is submitted all at once.

AA-WER Streaming consists of around 8 hours of audio from three datasets: AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions.

AA-WER Streaming Index is a dataset-weighted average consistent with our offline STT benchmark: AA-AgentTalk 50% / VoxPopuli 25% / Earnings22 25%. WER is audio-duration-weighted within each dataset; Time to Final and Time to First Partial are simple averages within each dataset, then dataset-weighted 50% / 25% / 25% overall.

Latenz

Zeit bis zur finalen Transkription

Sekunden bis zur finalen Transkription nach Sprachende

Starts at the SileroVAD-detected end of speech. For models that support forced endpointing, we send the endpoint request at this point and stop the timer on the next final transcript from the model. If the model already shared their last final before SileroVAD fired and no more finals arrive afterwards, we use that last final.

For models that do not support forced endpointing, we use the first natural final within 2 seconds of speech end. If no final arrives within 2 seconds, we use the first partial after 2 seconds, or the latest partial before 2 seconds if nothing arrives after. WER is computed on the combined previous finals plus the selected final or partial transcript.

Zeit bis zur ersten partiellen Transkription nach Sprachende

Sekunden bis zur ersten partiellen Transkription nach Sprachende

Starts at the SileroVAD-detected end of speech and stops on the first transcript-bearing event after speech end, whether partial or final. If no transcript arrives after speech end, we use the latest transcript before speech end as the fallback snapshot.

Preis

Preis der Transkription

USD per 1000 minutes of audio

Estimated cost in USD to transcribe 1,000 minutes of audio, normalized across providers with different billing models, and including billed reasoning tokens where available. Further detail on the methodology page.