语音转文本 AI 模型与服务商排行榜

比较不同语音转文本模型和服务商的词错误率、速度与价格。

有关更多详细信息,请参阅我们的方法论页面

亮点

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)
Seconds to final transcript after speech end · weighted average of samples in AA-WER Streaming · Lower is better
USD per 1000 minutes of audio · Lower is better

AA-WER Streaming 指数 vs. 最终转写时间

AA-WER Streaming 指数 vs. 最终转写时间

检测到语音结束后最终转写中错误转写的词占比 vs. 语音结束后到最终转写的秒数
Most attractive quadrant
Pareto line

AA-WER Streaming 指数 - 最终转写

检测到语音结束后最终转写中错误转写的词占比

AA-WER Streaming - 最终转写:AA-AgentTalk 数据集

在 AA-AgentTalk 数据集上,检测到语音结束后最终转写中错误转写的词占比;越低越好

AA-WER Streaming 指数(首次部分)vs. 语音结束后首次部分转写时间

AA-WER Streaming 指数(首次部分)vs. 语音结束后首次部分转写时间

检测到语音结束后首次部分转写中错误转写的词占比 vs. 语音结束后到首次部分转写的秒数
Most attractive quadrant
Pareto line

语音结束后首次部分转写的 AA-WER Streaming 指数

检测到语音结束后首次部分转写中错误转写的词占比

语音结束后首次部分转写的 AA-WER Streaming:AA-AgentTalk 数据集

在 AA-AgentTalk 数据集上,检测到语音结束后首次部分转写中错误转写的词占比;越低越好

AA-WER Streaming - 最终转写与语音结束后首次部分转写对比

% of words transcribed incorrectly · Lower is better · AA-WER Streaming incorporates 3 datasets: AA-AgentTalk (50%), VoxPopuli (25%), Earnings22 (25%)

延迟

最终转写时间

语音结束后到最终转写的秒数

语音结束后首次部分转写时间

语音结束后到首次部分转写的秒数

价格

转写价格

USD per 1000 minutes of audio