benchmarks/speech-to-text by coval
live independent STT benchmark · latency & WER

Independent Live Speech-to-Text (STT/ASR) Benchmark

Why did we build this benchmark?

People pick speech-to-text on two questions: which API is fastest, and which has the lowest Word Error Rate. This board measures those axes on identical audio. They are not one ranking, so we do not crown a single winner. We show the Pareto frontier first: the models that are not beaten on both speed and accuracy at once. Right now that is 4 of 24 models with both metrics; the rest lose on latency and WER together. The fastest 11 models on median TTFS span 79 ms, which is usually smaller than endpointing, network, and the rest of a voice pipeline. It is live because that set can change: models enter the Pareto frontier as the rolling 7-day window moves. A low WER still does not mean important fields such as emails, URLs, or IDs were captured correctly for a downstream task. For that, use the Structured Speech-to-Text Benchmark.

Definition

Headline metrics: TTFS (Time to Final Segment, the lag after speech ends) and WER (share of words wrong). Same pinned dataset, production-realistic conditions, mirrored here daily with attribution. Read p95 TTFS with average WER: the latency tail is what a realtime pipeline feels. Last 7d.

Pareto frontier: Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.

How to pick from this board

  • Live / realtime transcription

    Do not rank by median TTFS. Start from the Pareto frontier, then read p95 TTFS and WER together. Median gaps of tens of milliseconds are usually smaller than the rest of a voice pipeline, and the lowest-median model can have much worse WER.

  • Human-read transcripts, captions, or search

    Rank by Word Error Rate. That is the right axis when a person reads the text.

  • Emails, phones, paths, or IDs must be exact

    Do not pick from this board. Go to the Structured Speech-to-Text Benchmark.

Is there a best STT model?

No, not as a single ranking. Fastest and lowest WER are often different systems, and ranking by median TTFS hides models that are beaten on both axes at once. That is why we show the Pareto frontier. Structured capture (emails, phones, IDs) is a different failure mode on a different board. For a combined buying guide, see the best speech-to-text models by use case.

Which speech-to-text API is best?

There is no single “best”. Lowest median TTFS is stt-rt-v5 (Soniox) at 49 ms with 5.7% WER; lowest WER is universal-3.5-pro (AssemblyAI) at 3.3%. Median TTFS is the wrong single ranking for realtime. Route by the constraint that actually binds:

Your constraintPick byCurrent leader
Live / realtime transcription: the transcript has to land, but tens of milliseconds of median TTFS rarely decide the pipelinep95 TTFS and WER togetherthe Pareto frontier (4 models) not the lowest median TTFS alone
Human-read transcripts: captions, notes, search, or analyticslowest WERuniversal-3.5-pro AssemblyAI · 3.3%
Emails, phones, paths, or IDs must be exactnot this boardStructured Speech-to-Text Benchmark

Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. For emails, phones, paths, and IDs, go to the Structured Speech-to-Text Benchmark. Full ranking below.

Run by Coval, a voice-AI evaluation platform

This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution. A public pinned set is a reproducibility target and a contamination risk: continuous re-running against a fixed set does not make that tension smaller over time.

synced from Covalrolling 7d aggregate · pulled 2026-08-25 06:31 UTC · this table a5be8e8f7ce3 · pinned dataset 0f717f72b55c (stable; does not move with sample counts) · methodology → · benchmarks.coval.ai → · open-source runner: re-run it yourself →

These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.

Speech-to-Text models ranked by Word Error Rate

Every model runs the same pinned dataset under the same conditions. Last 7d. The Pareto frontier below is the useful cut; the full ranking after it is every model, sorted by average WER.

The Pareto frontier

Same meaning as in the definition at the top: a model is here only if nothing else is both as fast or faster and as accurate or more accurate. Example: nova-3 and nova-2 are both about 95 ms median TTFS, but nova-3 is 6.3% WER against 8.0% for nova-2, so nova-2 is not in this table. Currently 4 of 24 models with both metrics. This is not a single winner: the fastest row can have much worse WER.

The Pareto frontier is not a fixed ranking. A week is long enough to change the set. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro. ink-2 and inworld-stt-1 joined between Aug 5 and Aug 10, after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms.

ModelTTFS medianTTFS p95WER avg
stt-rt-v5
Soniox
49 ms77 ms5.7%
inworld-stt-1
Inworld AI
63 ms89 ms5.1%
ink-2
Cartesia
105 ms138 ms4.8%
universal-3.5-pro
AssemblyAI
128 ms317 ms3.3%

Full ranking, sorted by WER

Lower is better. Median TTFS is shown for context.

#ModelTTFS medianTTFS p95WER avgSamples
1universal-3.5-pro
AssemblyAI
128 ms317 ms3.3%6,787
2realtime
Reson8
266 ms300 ms3.5%6,782
3chirp_3
Google
780 ms1,001 ms4.2%6,788
4enhanced
Speechmatics
324 ms512 ms4.3%6,785
5gpt-4o-transcribe
OpenAI
704 ms954 ms4.6%6,786
6grok-stt
xAI
194 ms263 ms4.7%6,787
7ink-2
Cartesia
105 ms138 ms4.8%6,781
8velma-2-stt-streaming-english-v2
Modulate
123 ms177 ms4.9%2,618
9chirp_2
Google
784 ms1,286 ms4.9%6,752
10gpt-4o-mini-transcribe
OpenAI
621 ms942 ms5.0%6,786
11inworld-stt-1
Inworld AI
63 ms89 ms5.1%6,775
12gpt-realtime-whisper
OpenAI
557 ms673 ms5.2%6,777
13default
Speechmatics
217 ms273 ms5.4%6,787
14pulse
Smallest
194 ms292 ms5.4%6,786
15scribe_v2_realtime
ElevenLabs
111 ms195 ms5.6%6,785
16stt-rt-v5
Soniox
49 ms77 ms5.7%6,787
17voxtral-mini-transcribe-realtime-2602
Mistral
325 ms536 ms6.1%6,762
18nova-3
Deepgram
95 ms141 ms6.3%6,780
19velma-2-stt-streaming
Modulate
92 ms1,028 ms6.5%2,604
20flux-general-en
Deepgram
----6.5%6,787
21universal-streaming
AssemblyAI
----7.0%6,771
22solaria-1
Gladia
398 ms2,574 ms7.3%6,767
23nova-2
Deepgram
96 ms155 ms8.0%6,777
24whisper-large-v3
Together AI
122 ms417 ms8.2%6,770
25flux-general-multi
Deepgram
----8.5%6,782
26default
Gradium
248 ms326 ms10.1%6,780
27parakeet-tdt-0.6b-v3
Together AI
62 ms137 ms11.2%6,764
28universal-streaming-multilingual
AssemblyAI
----11.8%2,589
29nemotron-3.5-asr-streaming-0.6b
Together AI
----15.5%6,737

3 rows have far fewer samples than the rest of the board (typical 6,780): velma-2-stt-streaming (2,604); velma-2-stt-streaming-english-v2 (2,618); universal-streaming-multilingual (2,589). Treat the whole row: WER and TTFS, including p95, as thinner evidence. A p95 on 2,589 samples rests on about 129 tail points. 5 models have no TTFS in this window. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →

What this snapshot does not tell you

  • Who decides that speech ended

    TTFS is lag after speech ends. If the provider's own VAD makes that call, aggressive endpointing can buy latency by clipping words. The published definition does not rule that out.

  • Text normalization and call region

    How numbers, casing, punctuation, and contractions are handled can move WER by whole points, and is not specified here. At 65–120 ms, which region the runner calls from is a large fraction of the number, and is also unstated.

  • What decides real deployments

    Hallucination on silence, endpointing behavior, timestamps, diarization, long-form drift, cost, and rate limits under load are not on this board. Emails, phones, paths, and IDs are on the Structured Speech-to-Text Benchmark.

Compare two speech-to-text providers directly

Provider vs provider on the same measured data: best model per axis, winner per axis:

Deepgram vs AssemblyAI · OpenAI vs Deepgram · Soniox vs Deepgram

Speech-to-Text benchmark: common questions

Which speech-to-text API is the fastest?

stt-rt-v5 (Soniox) has the lowest median Time to Final Segment in the current 7d window: 49 ms. That is not a pick by itself: its WER is 5.7% against 3.3% for the WER leader. The fastest 11 models on median TTFS span 79 ms. Median gaps of tens of milliseconds are usually smaller than endpointing, network, and the rest of a voice pipeline. For realtime, read p95 TTFS and WER together; start from the Pareto frontier.

Which speech-to-text API is the most accurate?

universal-3.5-pro (AssemblyAI) has the lowest average Word Error Rate: 3.3%. Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.

Is OpenAI (GPT-4o Transcribe) the best STT API?

On the measured axes: OpenAI's fastest model, gpt-realtime-whisper, ranks #20 of 24 on median TTFS (557 ms), and its most accurate, gpt-4o-transcribe, ranks #5 of 29 on WER (4.6%). Lowest median TTFS is stt-rt-v5 (49 ms); lowest WER is universal-3.5-pro (3.3%). Median TTFS is the wrong single ranking for realtime. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.

What is Time to Final Segment (TTFS)?

Time from the end of speech to the final transcript segment being returned. The lag before a downstream task can use what was said. Lower is better. The tail (p95) is what a realtime pipeline feels; median gaps of tens of milliseconds are usually smaller than endpointing, network, and everything after the transcript.

What is Word Error Rate (WER)?

Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.

Does this benchmark measure price?

No. Pricing is not measured. The benchmark measures latency (TTFS) and Word Error Rate on identical audio. Hallucination on silence, endpointing, timestamps, diarization, long-form drift, and rate limits are also unmeasured; check each provider's pricing page for cost.

Is this STT benchmark independent? Who runs it?

Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.

Why is this a live speech-to-text benchmark?

Because the ranking is a rolling 7-day window. Coval re-runs continuously; this page mirrors that window. Models move on latency and WER, and that is enough for new systems to enter the Pareto frontier. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro on the frontier. Between Aug 5 and Aug 10, ink-2 and inworld-stt-1 entered after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms. A static leaderboard would still be showing last week's set. Last pull 2026-08-25 06:31 UTC. The table hash on this page is of the figures shown; a Coval run ID does not identify this window.

Does a low Word Error Rate mean important fields survive for downstream tasks?

No. WER treats every word equally, so a transcript can look clean and still get an email, URL, or ID wrong. This board measures raw Word Error Rate and Time to Final Segment only. If those fields have to be correct for a downstream task, go to the Structured Speech-to-Text Benchmark.

What is the Pareto frontier on this board?

Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.

Why do some models have no TTFS?

5 of 29 models in this window have no Time to Final Segment. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Treat missing latency as undisclosed, not as zero.

Are multilingual and English-only models compared fairly?

They share this board. universal-streaming-multilingual, flux-general-multi sit next to English-only rows. The pinned set is not characterized by language, accent, or domain here. If it is English-dominant, multilingual models can look worse on WER for a reason that is not their English quality. Treat mixed-language rows as a different category.