Cartesia vs Deepgram: Text-to-Speechlatency & accuracy
Measured head-to-head, not marketing: the verdict is split: Cartesia is faster on median TTFA (271 ms vs 292 ms), while Deepgram is more accurate (5.4% WER vs 5.9%). Independent data, refreshed daily, on identical inputs. Voice quality and price are not measured.
Cartesia or Deepgram: which is better?
Each provider's best model on each measured axis, last 7d. Lower is better on both:
| Axis | Cartesia (best model) | Deepgram (best model) | Measured winner |
|---|---|---|---|
| Speed: median TTFA | sonic-3.5 · 271 ms | aura-2-thalia-en · 292 ms | Cartesia |
| Accuracy: avg WER | sonic-3.5 · 5.9% | aura-2-thalia-en · 5.4% | Deepgram |
Not measured here: voice quality / naturalness and price. Treat those as vendor claims until measured. Full leaderboard context: the independent text-to-speech benchmark →
Every Cartesia and Deepgram model, measured
All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFA), last 7d:
| Model | Provider | TTFA median | TTFA p95 | WER | Samples |
|---|---|---|---|---|---|
| sonic-3.5 | Cartesia | 271 ms | 352 ms | 5.9% | 3,359 |
| aura-2-thalia-en | Deepgram | 292 ms | 541 ms | 5.4% | 3,352 |
Data by Coval
Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.
Cartesia vs Deepgram: common questions
Is Cartesia faster than Deepgram for text-to-speech?
On the current 7d window, Cartesia's fastest model (sonic-3.5) has a median TTFA of 271 ms, vs 292 ms for Deepgram's fastest (aura-2-thalia-en), so Cartesia is faster on measured median latency. Distributions matter too: p95 is 352 ms for Cartesia vs 541 ms for Deepgram.
Which is more accurate, Cartesia or Deepgram?
By Word Error Rate: Cartesia's best model (sonic-3.5) averages 5.9%, vs 5.4% for Deepgram's best (aura-2-thalia-en). Deepgram leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.
Which should I pick, Cartesia or Deepgram?
The measured verdict is split: Cartesia is faster (median TTFA), Deepgram is more accurate (WER). Pick by your constraint: realtime workloads care about latency; fidelity-critical workloads care about WER. Note voice quality / naturalness and price are not measured here.
Where does this data come from?
Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).