benchmarks/text-to-speech by coval
voice AI benchmark · by Coval

Independent Text-to-Speech (TTS) benchmark

An independent benchmark of text-to-speech providers on Time to First Audio (TTFA) and Word Error Rate (WER), measured under production-realistic conditions by Coval. Fastest is vui (101 ms median TTFA), lowest word error rate is tts-rt-v1 (3.7%). Mirrored here with attribution; last 7d.

Which text-to-speech API is best?

There is no single “best”. The measured split: vui (Fluxions) is the fastest at 101 ms median TTFA, while tts-rt-v1 (Soniox) is the most accurate at 3.7% WER. Route by the axis your workload is constrained by:

Your constraintPick byCurrent leader
Real-time voice agent — the reply must start speaking instantlylowest median TTFAvui Fluxions · 101 ms
Content fidelity — names, numbers, and scripts must come out exactly rightlowest WERtts-rt-v1 Soniox · 3.7%

Not measured here: voice quality / naturalness and price: treat those as vendor claims until measured. Full ranking below.

Run by Coval, a voice-AI evaluation platform

This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution.

synced from Covalrolling 7d aggregate · pulled 2026-08-25 06:31 UTC · this table f758d5ed8a45 · pinned dataset 0f717f72b55c (stable; does not move with sample counts) · methodology → · benchmarks.coval.ai → · open-source runner: re-run it yourself →

These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.

Text-to-Speech models ranked by Time to First Audio

Every model runs the same pinned dataset under the same conditions. TTFA (median) is the headline latency; WER is accuracy (lower is better on both). Ranked fastest-first, last 7d:

#ModelTTFA medianTTFA p95WERSamples
1vui
Fluxions
101 ms162 ms6.2%2,304
2palabra-tts-v1
Palabra
105 ms154 ms6.0%3,226
3inworld-tts-2-flash
Inworld AI
113 ms174 ms4.8%3,350
4inworld-tts-2
Inworld AI
162 ms243 ms4.6%3,351
5blizzard
LMNT
217 ms405 ms7.7%3,350
6tts-rt-v2
Soniox
228 ms275 ms3.9%3,349
7tts-rt-v1
Soniox
229 ms283 ms3.7%3,350
8mistv3
Rime
252 ms284 ms5.9%3,351
9sonic-3.5
Cartesia
272 ms353 ms5.9%3,351
10aura-2-thalia-en
Deepgram
286 ms539 ms5.4%3,344
11coda
Rime
312 ms404 ms5.1%3,351
12s2.1-pro
Fish Audio
316 ms513 ms4.7%3,351
13lightning_v3.1_pro
Smallest
331 ms614 ms4.4%3,339
14s2.1-pro-free
Fish Audio
350 ms1,331 ms4.6%3,317
15grok-tts
xAI
383 ms530 ms4.9%3,347
16default
Gradium
386 ms485 ms4.6%3,350
17s1
Fish Audio
387 ms545 ms4.9%3,351
18simba-3.2
Speechify
404 ms654 ms4.3%3,348
19simba-3.0
Speechify
414 ms585 ms5.2%3,351
20speech-2.8-turbo
Minimax
426 ms507 ms4.2%258
21speech-2.8-hd
Minimax
471 ms588 ms3.9%258
22chirp-3-hd
Google
493 ms887 ms5.1%3,350
23falcon-2
Murf
557 ms773 ms5.4%3,351
24qwen3-tts-flash-realtime
Alibaba
697 ms816 ms8.8%3,351
25gpt-4o-mini-tts
OpenAI
775 ms3,907 ms4.8%3,350

Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →

Compare two text-to-speech providers directly

Provider vs provider on the same measured data: best model per axis, winner per axis:

ElevenLabs vs Cartesia · ElevenLabs vs OpenAI · Cartesia vs Deepgram

Text-to-Speech benchmark: common questions

Which text-to-speech API is the fastest?

vui (Fluxions) has the lowest median Time to First Audio in the current 7d window: 101 ms. Next is palabra-tts-v1 (Palabra) at 105 ms. Latency is TTFA: time from sending the synthesis request to the first byte of audio coming back — the pause a caller hears before the voice starts speaking. Lower is better.

Which text-to-speech API is the most accurate?

tts-rt-v1 (Soniox) has the lowest average Word Error Rate: 3.7%. Share of words wrong in the synthesized audio: the audio is transcribed back with a reference speech-to-text model and compared to the input text, which catches skipped, garbled, and hallucinated words. Lower is better.

What is Time to First Audio (TTFA)?

Time from sending the synthesis request to the first byte of audio coming back — the pause a caller hears before the voice starts speaking. Lower is better.

What is Word Error Rate (WER)?

Share of words wrong in the synthesized audio: the audio is transcribed back with a reference speech-to-text model and compared to the input text, which catches skipped, garbled, and hallucinated words. Lower is better.

Does this benchmark measure voice quality or price?

No. It measures latency (TTFA) and Word Error Rate on identical inputs. Voice quality / naturalness (how human the voice sounds) and price are not measured; treat those as vendor claims until measured.

Is this TTS benchmark independent? Who runs it?

Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.

How fresh is this data?

The table is a rolling 7d aggregate, not a single Coval job. Coval re-runs continuously (roughly every 30 minutes); this page mirrors the window and shows a hash of the figures on this page. Last pull 2026-08-25 06:31 UTC. A Coval run ID is not printed next to the table because it does not identify these numbers.