最良のプロバイダー音声テキスト読み上げ(TTS)モデル

音声生成モデルの品質、速度、料金を、各モデル独自の音声が利用できる場合はそれを使用して比較します。

詳細については、方法論ページをご覧ください。

品質

プロバイダー音声アリーナの品質Elo

Arena Elo: average Elo rating of the model · Higher is better

Relative Elo score of the models as determined by responses from users in Artificial Analysis' Speech Arena. Some models may not be shown due to not yet having enough votes.

料金

料金

Price: USD per 1M characters of text · Lower is better

Price per 1M characters of text. For detail on how we calculate price for providers which price based on inference time or subscription plans, see our methodology page.

速度係数

1秒あたりの文字数

Characters processed per second: # of characters per second of generation time · Higher is better

Number of characters processed per second of generation time. Higher values indicate faster generation speeds.

よくある質問

Sonic 3.6は現在、Artificial Analysisのテキスト読み上げモデル比較で最高品質を持ち、品質Eloは1282です。 Sonic 3.6を見る

Polly Standardは現在、代表的なAPIエンドポイントでの比較において最速のテキスト読み上げモデルであり、Amazon Bedrockは1秒あたり1,315.0文字を配信します。 Polly Standardのプロバイダーを見る

Kokoro 82M v1.0は現在、代表的なAPIエンドポイントでの比較において最も安価なテキスト読み上げモデルであり、Replicateの価格は100万文字あたり$0.65です。 Kokoro 82M v1.0のプロバイダーを見る

Sonic 3.6(Cartesia経由)、Realtime TTS-2(Inworld経由)、Simba 3.2(SpeechifyAI経由)は現在、比較の中で品質対価格のトレードオフが特に強いエンドポイントの一部です。これらの代表的なエンドポイントは現在の品質対価格のフロンティアに位置しており、同時により安くて品質が高い代替はほとんどありません。 Sonic 3.6のプロバイダーを見る