benchmarks/tts by coval/cartesia vs deepgram
text-to-speech head-to-head · data by Coval

Cartesia vs Deepgram: Text-to-Speechlatency & accuracy

Measured head-to-head, not marketing: the verdict is split: Cartesia is faster on median TTFA (271 ms vs 292 ms), while Deepgram is more accurate (5.4% WER vs 5.9%). Independent data, refreshed daily, on identical inputs. Voice quality and price are not measured.

Cartesia or Deepgram: which is better?

Each provider's best model on each measured axis, last 7d. Lower is better on both:

AxisCartesia (best model)Deepgram (best model)Measured winner
Speed: median TTFAsonic-3.5 · 271 msaura-2-thalia-en · 292 msCartesia
Accuracy: avg WERsonic-3.5 · 5.9%aura-2-thalia-en · 5.4%Deepgram

Not measured here: voice quality / naturalness and price. Treat those as vendor claims until measured. Full leaderboard context: the independent text-to-speech benchmark →

Every Cartesia and Deepgram model, measured

All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFA), last 7d:

ModelProviderTTFA medianTTFA p95WERSamples
sonic-3.5Cartesia271 ms352 ms5.9%3,359
aura-2-thalia-enDeepgram292 ms541 ms5.4%3,352

Data by Coval

Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.

synced from Coval2026-08-27 17:09 UTC · full TTS benchmark → · methodology →

Cartesia vs Deepgram: common questions

Is Cartesia faster than Deepgram for text-to-speech?

On the current 7d window, Cartesia's fastest model (sonic-3.5) has a median TTFA of 271 ms, vs 292 ms for Deepgram's fastest (aura-2-thalia-en), so Cartesia is faster on measured median latency. Distributions matter too: p95 is 352 ms for Cartesia vs 541 ms for Deepgram.

Which is more accurate, Cartesia or Deepgram?

By Word Error Rate: Cartesia's best model (sonic-3.5) averages 5.9%, vs 5.4% for Deepgram's best (aura-2-thalia-en). Deepgram leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.

Which should I pick, Cartesia or Deepgram?

The measured verdict is split: Cartesia is faster (median TTFA), Deepgram is more accurate (WER). Pick by your constraint: realtime workloads care about latency; fidelity-critical workloads care about WER. Note voice quality / naturalness and price are not measured here.

Where does this data come from?

Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).