OpenAI vs Deepgram: Speech-to-Textlatency & accuracy
Measured head-to-head, not marketing: the verdict is split: Deepgram is faster on median TTFS (96 ms vs 557 ms), while OpenAI is more accurate (4.4% WER vs 6.0%). Median TTFS gaps of tens of milliseconds are usually not a pick reason; read p95 TTFS with WER. Independent data, refreshed daily, on identical inputs. Price and whether emails, phones, paths, and IDs survive verbatim are not measured.
OpenAI or Deepgram: which is better?
Each provider's best model on each measured axis, last 7d. Lower is better on both:
| Axis | OpenAI (best model) | Deepgram (best model) | Measured winner |
|---|---|---|---|
| Speed: median TTFS | gpt-realtime-whisper · 557 ms | nova-3 · 96 ms | Deepgram |
| Accuracy: avg WER | gpt-4o-transcribe · 4.4% | nova-3 · 6.0% | OpenAI |
Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. Treat those as vendor claims until measured. Full leaderboard context: the independent speech-to-text benchmark →
Every OpenAI and Deepgram model, measured
All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFS), last 7d:
| Model | Provider | TTFS median | TTFS p95 | WER | Samples |
|---|---|---|---|---|---|
| nova-3 | Deepgram | 96 ms | 140 ms | 6.0% | 6,736 |
| nova-2 | Deepgram | 96 ms | 154 ms | 7.9% | 6,712 |
| gpt-realtime-whisper | OpenAI | 557 ms | 675 ms | 4.9% | 6,733 |
| gpt-4o-mini-transcribe | OpenAI | 649 ms | 976 ms | 4.8% | 6,738 |
| gpt-4o-transcribe | OpenAI | 711 ms | 973 ms | 4.4% | 6,739 |
Data by Coval
Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.
OpenAI vs Deepgram: common questions
Is OpenAI faster than Deepgram for speech-to-text?
On the current 7d window, OpenAI's fastest model (gpt-realtime-whisper) has a median TTFS of 557 ms, vs 96 ms for Deepgram's fastest (nova-3), so Deepgram is faster on measured median latency. Distributions matter too: p95 is 675 ms for OpenAI vs 140 ms for Deepgram.
Which is more accurate, OpenAI or Deepgram?
By Word Error Rate: OpenAI's best model (gpt-4o-transcribe) averages 4.4%, vs 6.0% for Deepgram's best (nova-3). OpenAI leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.
Which should I pick, OpenAI or Deepgram?
The measured verdict is split: Deepgram is faster on median TTFS, OpenAI is more accurate (WER). A few milliseconds of median TTFS is usually not a pick reason; read p95 TTFS and WER together. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.
Where does this data come from?
Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).