Best Provider Voice Text to Speech (TTS) Models
Compare quality, speed, and price of speech generation models with every model speaking its own native voices where available.
For further details, see our methodology page.
QualityUpdated
Provider Voice Arena Preference Elo
Pronunciation Robustness
Pronunciation Robustness by Category
Pricing
Price
Throughput
Characters Per Second
Frequently Asked Questions
Eleven v4 currently has the highest quality in the Artificial Analysis Text to Speech models comparison, with a Quality Elo of 1315. View Eleven v4
Polly Standard is currently the fastest Text to Speech model in the comparison at representative API endpoints, with Amazon Bedrock delivering 1,148.8 characters per second. View Polly Standard providers
Kokoro 82M v1.0 is currently the cheapest Text to Speech model in the comparison at representative API endpoints, with Replicate priced at $0.65 per 1 million characters. View Kokoro 82M v1.0 providers
Eleven v4 via ElevenLabs, Sonic 3.6 via Cartesia, and Gemini 3.8 Flash TTS via Google currently offer some of the strongest quality-for-price tradeoffs in the comparison. These representative endpoints sit on the current quality-versus-price frontier, meaning few alternatives are both cheaper and higher quality at the same time. View Eleven v4 providers