Independent Text-to-Speech (TTS) benchmark
An independent benchmark of text-to-speech providers on Time to First Audio (TTFA) and Word Error Rate (WER), measured under production-realistic conditions by Coval. Fastest is vui (40 ms median TTFA), lowest word error rate is qwen3-tts-1.7b (1.5%). Mirrored here with attribution; last 7d.
Which text-to-speech API is best?
There is no single “best”. The measured split: vui (Fluxions) is the fastest at 40 ms median TTFA, while qwen3-tts-1.7b (Baseten) is the most accurate at 1.5% WER. Route by the axis your workload is constrained by:
| Your constraint | Pick by | Current leader |
|---|---|---|
| Real-time voice agent — the reply must start speaking instantly | lowest median TTFA | vui Fluxions · 40 ms |
| Content fidelity — names, numbers, and scripts must come out exactly right | lowest WER | qwen3-tts-1.7b Baseten · 1.5% |
Not measured here: voice quality / naturalness and price: treat those as vendor claims until measured. Full ranking below.
Run by Coval, a voice-AI evaluation platform
This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution.
These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.
Text-to-Speech models ranked by Time to First Audio
Every model runs the same pinned dataset under the same conditions. TTFA (median) is the headline latency; WER is accuracy (lower is better on both). Ranked fastest-first, last 7d:
| # | Model | TTFA median | TTFA p95 | WER | Samples |
|---|---|---|---|---|---|
| 1 | vui Fluxions | 40 ms | 70 ms | 2.0% | 592 |
| 2 | gradium-tts-beta-202609 Gradium | 54 ms | 102 ms | 1.9% | 670 |
| 3 | inworld-tts-2-flash Inworld AI | 62 ms | 93 ms | 1.6% | 670 |
| 4 | qwen3-tts-fast Nari | 62 ms | 99 ms | 1.9% | 670 |
| 5 | simba-3.0 Speechify | 98 ms | 147 ms | 1.6% | 670 |
| 6 | qwen3-tts-1.7b Baseten | 104 ms | 140 ms | 1.5% | 672 |
| 7 | palabra-tts-v1 Palabra | 104 ms | 208 ms | 2.9% | 544 |
| 8 | simba-3.2 Speechify | 107 ms | 171 ms | 2.0% | 669 |
| 9 | inworld-tts-2 Inworld AI | 149 ms | 192 ms | 1.6% | 670 |
| 10 | eleven_flash_v2_5 ElevenLabs | 187 ms | 230 ms | 1.8% | 663 |
| 11 | eleven_v4_turbo ElevenLabs | 194 ms | 280 ms | 1.7% | 670 |
| 12 | flux-haley-en Deepgram | 200 ms | 273 ms | 1.7% | 638 |
| 13 | tts-rt-v1 Soniox | 254 ms | 313 ms | 2.7% | 557 |
| 14 | aura-2-thalia-en Deepgram | 260 ms | 517 ms | 1.8% | 670 |
| 15 | tts-rt-v2 Soniox | 262 ms | 319 ms | 2.8% | 553 |
| 16 | dd-etts-3.3 Deepdub | 270 ms | 462 ms | 1.9% | 670 |
| 17 | falcon-2 Murf | 274 ms | 377 ms | 1.8% | 669 |
| 18 | sonic-3.5 Cartesia | 284 ms | 434 ms | 1.7% | 670 |
| 19 | sonic-3.6 Cartesia | 294 ms | 385 ms | 1.6% | 670 |
| 20 | s2.1-pro Fish Audio | 305 ms | 668 ms | 1.6% | 670 |
| 21 | eleven_v3_conversational ElevenLabs | 320 ms | 398 ms | 1.8% | 663 |
| 22 | lightning_v3.1_pro Smallest | 357 ms | 664 ms | 1.6% | 670 |
| 23 | grok-tts xAI | 358 ms | 507 ms | 1.6% | 668 |
| 24 | s1 Fish Audio | 381 ms | 510 ms | 1.8% | 670 |
| 25 | aura-2-en Cloudflare | 532 ms | 1,292 ms | 2.1% | 670 |
| 26 | chirp-3-hd | 533 ms | 949 ms | 1.6% | 670 |
| 27 | qwen3-tts-flash-realtime Alibaba | 731 ms | 790 ms | 2.9% | 670 |
| 28 | s2.1-pro-free Fish Audio | 856 ms | 1,116 ms | 1.6% | 666 |
| 29 | gpt-4o-mini-tts OpenAI | 2,380 ms | 5,356 ms | 1.9% | 662 |
Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →
Compare two text-to-speech providers directly
Provider vs provider on the same measured data: best model per axis, winner per axis:
ElevenLabs vs Cartesia → · ElevenLabs vs OpenAI → · Cartesia vs Deepgram →
Text-to-Speech benchmark: common questions
Which text-to-speech API is the fastest?
vui (Fluxions) has the lowest median Time to First Audio in the current 7d window: 40 ms. Next is gradium-tts-beta-202609 (Gradium) at 54 ms. Latency is TTFA: time from sending the synthesis request to the first byte of audio coming back — the pause a caller hears before the voice starts speaking. Lower is better.
Which text-to-speech API is the most accurate?
qwen3-tts-1.7b (Baseten) has the lowest average Word Error Rate: 1.5%. Share of words wrong in the synthesized audio: the audio is transcribed back with a reference speech-to-text model and compared to the input text, which catches skipped, garbled, and hallucinated words. Lower is better.
Is ElevenLabs the best TTS API?
On the measured axes: ElevenLabs's fastest model, eleven_flash_v2_5, ranks #10 of 29 on median TTFA (187 ms), and its most accurate, eleven_v4_turbo, ranks #13 of 29 on WER (1.7%). The current leaders are vui (40 ms) and qwen3-tts-1.7b (1.5% WER). "Best" depends on which axis your workload is constrained by, and voice quality / naturalness and price are not measured here.
What is Time to First Audio (TTFA)?
Time from sending the synthesis request to the first byte of audio coming back — the pause a caller hears before the voice starts speaking. Lower is better.
What is Word Error Rate (WER)?
Share of words wrong in the synthesized audio: the audio is transcribed back with a reference speech-to-text model and compared to the input text, which catches skipped, garbled, and hallucinated words. Lower is better.
Does this benchmark measure voice quality or price?
No. It measures latency (TTFA) and Word Error Rate on identical inputs. Voice quality / naturalness (how human the voice sounds) and price are not measured; treat those as vendor claims until measured.
Is this TTS benchmark independent? Who runs it?
Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.
How fresh is this data?
The table is a rolling 7d aggregate, not a single Coval job. Coval re-runs continuously (roughly every 30 minutes); this page mirrors the window and shows a hash of the figures on this page. Last pull 2026-10-09 12:20 UTC. A Coval run ID is not printed next to the table because it does not identify these numbers.