Soniox vs Deepgram: Speech-to-Textlatency & accuracy
Measured head-to-head, not marketing: Soniox currently leads both measured axes: median TTFS and Word Error Rate. Independent data, refreshed daily, on identical inputs. Price and whether emails, phones, paths, and IDs survive verbatim are not measured.
Soniox or Deepgram: which is better?
Each provider's best model on each measured axis, last 7d. Lower is better on both:
| Axis | Soniox (best model) | Deepgram (best model) | Measured winner |
|---|---|---|---|
| Speed: median TTFS | stt-rt-v5 · 49 ms | nova-3 · 96 ms | Soniox |
| Accuracy: avg WER | stt-rt-v5 · 5.5% | nova-3 · 6.0% | Soniox |
Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. Treat those as vendor claims until measured. Full leaderboard context: the independent speech-to-text benchmark →
Every Soniox and Deepgram model, measured
All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFS), last 7d:
| Model | Provider | TTFS median | TTFS p95 | WER | Samples |
|---|---|---|---|---|---|
| stt-rt-v5 | Soniox | 49 ms | 79 ms | 5.5% | 6,746 |
| nova-3 | Deepgram | 96 ms | 140 ms | 6.0% | 6,736 |
| nova-2 | Deepgram | 96 ms | 154 ms | 7.9% | 6,712 |
Data by Coval
Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.
Soniox vs Deepgram: common questions
Is Soniox faster than Deepgram for speech-to-text?
On the current 7d window, Soniox's fastest model (stt-rt-v5) has a median TTFS of 49 ms, vs 96 ms for Deepgram's fastest (nova-3), so Soniox is faster on measured median latency. Distributions matter too: p95 is 79 ms for Soniox vs 140 ms for Deepgram.
Which is more accurate, Soniox or Deepgram?
By Word Error Rate: Soniox's best model (stt-rt-v5) averages 5.5%, vs 6.0% for Deepgram's best (nova-3). Soniox leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.
Which should I pick, Soniox or Deepgram?
Soniox currently leads both measured axes (median TTFS and WER). Full per-model distributions are in the table below; price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.
Where does this data come from?
Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).