Deepgram vs AssemblyAI: Speech-to-Textlatency & accuracy
Measured head-to-head, not marketing: the verdict is split: Deepgram is faster on median TTFS (96 ms vs 140 ms), while AssemblyAI is more accurate (3.1% WER vs 6.0%). Median TTFS gaps of tens of milliseconds are usually not a pick reason; read p95 TTFS with WER. Independent data, refreshed daily, on identical inputs. Price and whether emails, phones, paths, and IDs survive verbatim are not measured.
Deepgram or AssemblyAI: which is better?
Each provider's best model on each measured axis, last 7d. Lower is better on both:
| Axis | Deepgram (best model) | AssemblyAI (best model) | Measured winner |
|---|---|---|---|
| Speed: median TTFS | nova-3 · 96 ms | universal-3.5-pro · 140 ms | Deepgram |
| Accuracy: avg WER | nova-3 · 6.0% | universal-3.5-pro · 3.1% | AssemblyAI |
Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. Treat those as vendor claims until measured. Full leaderboard context: the independent speech-to-text benchmark →
Every Deepgram and AssemblyAI model, measured
All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFS), last 7d:
| Model | Provider | TTFS median | TTFS p95 | WER | Samples |
|---|---|---|---|---|---|
| nova-3 | Deepgram | 96 ms | 140 ms | 6.0% | 6,736 |
| nova-2 | Deepgram | 96 ms | 154 ms | 7.9% | 6,712 |
| universal-3.5-pro | AssemblyAI | 140 ms | 345 ms | 3.1% | 6,736 |
Data by Coval
Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.
Deepgram vs AssemblyAI: common questions
Is Deepgram faster than AssemblyAI for speech-to-text?
On the current 7d window, Deepgram's fastest model (nova-3) has a median TTFS of 96 ms, vs 140 ms for AssemblyAI's fastest (universal-3.5-pro), so Deepgram is faster on measured median latency. Distributions matter too: p95 is 140 ms for Deepgram vs 345 ms for AssemblyAI.
Which is more accurate, Deepgram or AssemblyAI?
By Word Error Rate: Deepgram's best model (nova-3) averages 6.0%, vs 3.1% for AssemblyAI's best (universal-3.5-pro). AssemblyAI leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.
Which should I pick, Deepgram or AssemblyAI?
The measured verdict is split: Deepgram is faster on median TTFS, AssemblyAI is more accurate (WER). A few milliseconds of median TTFS is usually not a pick reason; read p95 TTFS and WER together. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.
Where does this data come from?
Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).