benchmarks/stt by coval/soniox vs deepgram
speech-to-text head-to-head · data by Coval

Soniox vs Deepgram: Speech-to-Textlatency & accuracy

Measured head-to-head, not marketing: Soniox currently leads both measured axes: median TTFS and Word Error Rate. Independent data, refreshed daily, on identical inputs. Price and whether emails, phones, paths, and IDs survive verbatim are not measured.

Soniox or Deepgram: which is better?

Each provider's best model on each measured axis, last 7d. Lower is better on both:

AxisSoniox (best model)Deepgram (best model)Measured winner
Speed: median TTFSstt-rt-v5 · 49 msnova-3 · 96 msSoniox
Accuracy: avg WERstt-rt-v5 · 5.5%nova-3 · 6.0%Soniox

Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. Treat those as vendor claims until measured. Full leaderboard context: the independent speech-to-text benchmark →

Every Soniox and Deepgram model, measured

All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFS), last 7d:

ModelProviderTTFS medianTTFS p95WERSamples
stt-rt-v5Soniox49 ms79 ms5.5%6,746
nova-3Deepgram96 ms140 ms6.0%6,736
nova-2Deepgram96 ms154 ms7.9%6,712

Data by Coval

Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.

synced from Coval2026-08-27 17:09 UTC · full STT benchmark → · methodology →

Soniox vs Deepgram: common questions

Is Soniox faster than Deepgram for speech-to-text?

On the current 7d window, Soniox's fastest model (stt-rt-v5) has a median TTFS of 49 ms, vs 96 ms for Deepgram's fastest (nova-3), so Soniox is faster on measured median latency. Distributions matter too: p95 is 79 ms for Soniox vs 140 ms for Deepgram.

Which is more accurate, Soniox or Deepgram?

By Word Error Rate: Soniox's best model (stt-rt-v5) averages 5.5%, vs 6.0% for Deepgram's best (nova-3). Soniox leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.

Which should I pick, Soniox or Deepgram?

Soniox currently leads both measured axes (median TTFS and WER). Full per-model distributions are in the table below; price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.

Where does this data come from?

Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).