Best Provider Voice Text to Speech (TTS) Models

Compare quality, speed, and price of speech generation models with every model speaking its own native voices where available.

For further details, see our methodology page.

Quality

Provider Voice Arena Quality Elo

Arena Elo: average Elo rating of the model · Higher is better

Relative Elo score of the models as determined by responses from users in Artificial Analysis' Speech Arena. Some models may not be shown due to not yet having enough votes.

Pricing

Price

Price: USD per 1M characters of text · Lower is better

Price per 1M characters of text. For detail on how we calculate price for providers which price based on inference time or subscription plans, see our methodology page.

Speed Factor

Characters Per Second

Characters processed per second: # of characters per second of generation time · Higher is better

Number of characters processed per second of generation time. Higher values indicate faster generation speeds.

Frequently Asked Questions

Sonic 3.6 currently has the highest quality in the Artificial Analysis Text to Speech models comparison, with a Quality Elo of 1,285. View Sonic 3.6

Polly Standard is currently the fastest Text to Speech model in the comparison at representative API endpoints, with Amazon Bedrock delivering 1,141.9 characters per second. View Polly Standard providers

Kokoro 82M v1.0 is currently the cheapest Text to Speech model in the comparison at representative API endpoints, with Replicate priced at $0.65 per 1 million characters. View Kokoro 82M v1.0 providers

Sonic 3.6 via Cartesia, Qwen-Audio-3.0-TTS-Plus via Alibaba Cloud, and Simba 3.2 via SpeechifyAI currently offer some of the strongest quality-for-price tradeoffs in the comparison. These representative endpoints sit on the current quality-versus-price frontier, meaning few alternatives are both cheaper and higher quality at the same time. View Sonic 3.6 providers