benchmarks/text-to-speech by coval
voice AI benchmark · by Coval

Independent Text-to-Speech (TTS) benchmark

An independent benchmark of text-to-speech providers on Time to First Audio (TTFA) and Word Error Rate (WER), measured under production-realistic conditions by Coval. Fastest is vui (40 ms median TTFA), lowest word error rate is qwen3-tts-1.7b (1.5%). Mirrored here with attribution; last 7d.

Which text-to-speech API is best?

There is no single “best”. The measured split: vui (Fluxions) is the fastest at 40 ms median TTFA, while qwen3-tts-1.7b (Baseten) is the most accurate at 1.5% WER. Route by the axis your workload is constrained by:

Your constraintPick byCurrent leader
Real-time voice agent — the reply must start speaking instantlylowest median TTFAvui Fluxions · 40 ms
Content fidelity — names, numbers, and scripts must come out exactly rightlowest WERqwen3-tts-1.7b Baseten · 1.5%

Not measured here: voice quality / naturalness and price: treat those as vendor claims until measured. Full ranking below.

Run by Coval, a voice-AI evaluation platform

This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution.

synced from Covalrolling 7d aggregate · pulled 2026-10-09 12:20 UTC · this table 347dfc341d8b · pinned dataset 09d98bd6450c (stable; does not move with sample counts) · methodology → · benchmarks.coval.ai → · open-source runner: re-run it yourself →

These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.

Text-to-Speech models ranked by Time to First Audio

Every model runs the same pinned dataset under the same conditions. TTFA (median) is the headline latency; WER is accuracy (lower is better on both). Ranked fastest-first, last 7d:

Top 8 per metric
#ModelTTFA medianTTFA p95WERSamples
1vui
Fluxions
40 ms70 ms2.0%592
2gradium-tts-beta-202609
Gradium
54 ms102 ms1.9%670
3inworld-tts-2-flash
Inworld AI
62 ms93 ms1.6%670
4qwen3-tts-fast
Nari
62 ms99 ms1.9%670
5simba-3.0
Speechify
98 ms147 ms1.6%670
6qwen3-tts-1.7b
Baseten
104 ms140 ms1.5%672
7palabra-tts-v1
Palabra
104 ms208 ms2.9%544
8simba-3.2
Speechify
107 ms171 ms2.0%669
9inworld-tts-2
Inworld AI
149 ms192 ms1.6%670
10eleven_flash_v2_5
ElevenLabs
187 ms230 ms1.8%663
11eleven_v4_turbo
ElevenLabs
194 ms280 ms1.7%670
12flux-haley-en
Deepgram
200 ms273 ms1.7%638
13tts-rt-v1
Soniox
254 ms313 ms2.7%557
14aura-2-thalia-en
Deepgram
260 ms517 ms1.8%670
15tts-rt-v2
Soniox
262 ms319 ms2.8%553
16dd-etts-3.3
Deepdub
270 ms462 ms1.9%670
17falcon-2
Murf
274 ms377 ms1.8%669
18sonic-3.5
Cartesia
284 ms434 ms1.7%670
19sonic-3.6
Cartesia
294 ms385 ms1.6%670
20s2.1-pro
Fish Audio
305 ms668 ms1.6%670
21eleven_v3_conversational
ElevenLabs
320 ms398 ms1.8%663
22lightning_v3.1_pro
Smallest
357 ms664 ms1.6%670
23grok-tts
xAI
358 ms507 ms1.6%668
24s1
Fish Audio
381 ms510 ms1.8%670
25aura-2-en
Cloudflare
532 ms1,292 ms2.1%670
26chirp-3-hd
Google
533 ms949 ms1.6%670
27qwen3-tts-flash-realtime
Alibaba
731 ms790 ms2.9%670
28s2.1-pro-free
Fish Audio
856 ms1,116 ms1.6%666
29gpt-4o-mini-tts
OpenAI
2,380 ms5,356 ms1.9%662

Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →

Compare two text-to-speech providers directly

Provider vs provider on the same measured data: best model per axis, winner per axis:

ElevenLabs vs Cartesia → · ElevenLabs vs OpenAI → · Cartesia vs Deepgram →

Text-to-Speech benchmark: common questions

Which text-to-speech API is the fastest?

vui (Fluxions) has the lowest median Time to First Audio in the current 7d window: 40 ms. Next is gradium-tts-beta-202609 (Gradium) at 54 ms. Latency is TTFA: time from sending the synthesis request to the first byte of audio coming back — the pause a caller hears before the voice starts speaking. Lower is better.

Which text-to-speech API is the most accurate?

qwen3-tts-1.7b (Baseten) has the lowest average Word Error Rate: 1.5%. Share of words wrong in the synthesized audio: the audio is transcribed back with a reference speech-to-text model and compared to the input text, which catches skipped, garbled, and hallucinated words. Lower is better.

Is ElevenLabs the best TTS API?

On the measured axes: ElevenLabs's fastest model, eleven_flash_v2_5, ranks #10 of 29 on median TTFA (187 ms), and its most accurate, eleven_v4_turbo, ranks #13 of 29 on WER (1.7%). The current leaders are vui (40 ms) and qwen3-tts-1.7b (1.5% WER). "Best" depends on which axis your workload is constrained by, and voice quality / naturalness and price are not measured here.

What is Time to First Audio (TTFA)?

Time from sending the synthesis request to the first byte of audio coming back — the pause a caller hears before the voice starts speaking. Lower is better.

What is Word Error Rate (WER)?

Share of words wrong in the synthesized audio: the audio is transcribed back with a reference speech-to-text model and compared to the input text, which catches skipped, garbled, and hallucinated words. Lower is better.

Does this benchmark measure voice quality or price?

No. It measures latency (TTFA) and Word Error Rate on identical inputs. Voice quality / naturalness (how human the voice sounds) and price are not measured; treat those as vendor claims until measured.

Is this TTS benchmark independent? Who runs it?

Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.

How fresh is this data?

The table is a rolling 7d aggregate, not a single Coval job. Coval re-runs continuously (roughly every 30 minutes); this page mirrors the window and shows a hash of the figures on this page. Last pull 2026-10-09 12:20 UTC. A Coval run ID is not printed next to the table because it does not identify these numbers.