benchmarks/stt by coval/openai vs deepgram
speech-to-text head-to-head · data by Coval

OpenAI vs Deepgram: Speech-to-Textlatency & accuracy

Measured head-to-head, not marketing: the verdict is split: Deepgram is faster on median TTFS (96 ms vs 557 ms), while OpenAI is more accurate (4.4% WER vs 6.0%). Median TTFS gaps of tens of milliseconds are usually not a pick reason; read p95 TTFS with WER. Independent data, refreshed daily, on identical inputs. Price and whether emails, phones, paths, and IDs survive verbatim are not measured.

OpenAI or Deepgram: which is better?

Each provider's best model on each measured axis, last 7d. Lower is better on both:

AxisOpenAI (best model)Deepgram (best model)Measured winner
Speed: median TTFSgpt-realtime-whisper · 557 msnova-3 · 96 msDeepgram
Accuracy: avg WERgpt-4o-transcribe · 4.4%nova-3 · 6.0%OpenAI

Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. Treat those as vendor claims until measured. Full leaderboard context: the independent speech-to-text benchmark →

Every OpenAI and Deepgram model, measured

All benchmarked models from both providers on the same pinned dataset, ranked fastest-first (median TTFS), last 7d:

ModelProviderTTFS medianTTFS p95WERSamples
nova-3Deepgram96 ms140 ms6.0%6,736
nova-2Deepgram96 ms154 ms7.9%6,712
gpt-realtime-whisperOpenAI557 ms675 ms4.9%6,733
gpt-4o-mini-transcribeOpenAI649 ms976 ms4.8%6,738
gpt-4o-transcribeOpenAI711 ms973 ms4.4%6,739

Data by Coval

Coval sells voice-agent evaluation infrastructure. Numbers are a rolling 7d aggregate on a pinned dataset under production-realistic conditions. Openbenchmarks mirrors the results with attribution; methodology and runner are open-source.

synced from Coval2026-08-27 17:09 UTC · full STT benchmark → · methodology →

OpenAI vs Deepgram: common questions

Is OpenAI faster than Deepgram for speech-to-text?

On the current 7d window, OpenAI's fastest model (gpt-realtime-whisper) has a median TTFS of 557 ms, vs 96 ms for Deepgram's fastest (nova-3), so Deepgram is faster on measured median latency. Distributions matter too: p95 is 675 ms for OpenAI vs 140 ms for Deepgram.

Which is more accurate, OpenAI or Deepgram?

By Word Error Rate: OpenAI's best model (gpt-4o-transcribe) averages 4.4%, vs 6.0% for Deepgram's best (nova-3). OpenAI leads on measured accuracy. Lower is better; WER is measured on identical inputs under production-realistic conditions.

Which should I pick, OpenAI or Deepgram?

The measured verdict is split: Deepgram is faster on median TTFS, OpenAI is more accurate (WER). A few milliseconds of median TTFS is usually not a pick reason; read p95 TTFS and WER together. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.

Where does this data come from?

Coval sells voice-agent evaluation infrastructure. Every model runs the same pinned dataset under production-realistic conditions. The table is a rolling 7d aggregate, not a single job; Openbenchmarks mirrors it with attribution. Methodology and runner are open-source (Apache-2.0).