benchmarks/speech-to-text by coval
live independent STT benchmark · latency & WER

Independent Live Speech-to-Text (STT/ASR) Benchmark

Why did we build this benchmark?

People pick speech-to-text on two questions: which API is fastest, and which has the lowest Word Error Rate. This board measures those axes on identical audio. They are not one ranking, so we do not crown a single winner. We show the Pareto frontier first: the models that are not beaten on both speed and accuracy at once. Right now that is 3 of 28 models with both metrics; the rest lose on latency and WER together. The fastest 8 models on median TTFS span 80 ms, which is usually smaller than endpointing, network, and the rest of a voice pipeline. It is live because that set can change: models enter the Pareto frontier as the rolling 7-day window moves. A low WER still does not mean important fields such as emails, URLs, or IDs were captured correctly for a downstream task. For that, use the Structured Speech-to-Text Benchmark.

Definition

Headline metrics: TTFS (Time to Final Segment, the lag after speech ends) and WER (share of words wrong). Same pinned dataset, production-realistic conditions, mirrored here daily with attribution. Read p95 TTFS with average WER: the latency tail is what a realtime pipeline feels. Last 7d.

Pareto frontier: Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.

How to pick from this board

  • Live / realtime transcription

    Do not rank by median TTFS. Start from the Pareto frontier, then read p95 TTFS and WER together. Median gaps of tens of milliseconds are usually smaller than the rest of a voice pipeline, and the lowest-median model can have much worse WER.

  • Human-read transcripts, captions, or search

    Rank by Word Error Rate. That is the right axis when a person reads the text.

  • Emails, phones, paths, or IDs must be exact

    Do not pick from this board. Go to the Structured Speech-to-Text Benchmark.

Is there a best STT model?

No, not as a single ranking. Fastest and lowest WER are often different systems, and ranking by median TTFS hides models that are beaten on both axes at once. That is why we show the Pareto frontier. Structured capture (emails, phones, IDs) is a different failure mode on a different board. For a combined buying guide, see the best speech-to-text models by use case.

Which speech-to-text API is best?

There is no single “best”. Lowest median TTFS is qwen3-asr-1.7b (Baseten) at 21 ms with 3.0% WER; lowest WER is universal-3.6-pro (AssemblyAI) at 2.1%. Median TTFS is the wrong single ranking for realtime. Route by the constraint that actually binds:

Your constraintPick byCurrent leader
Live / realtime transcription: the transcript has to land, but tens of milliseconds of median TTFS rarely decide the pipelinep95 TTFS and WER togetherthe Pareto frontier (3 models) not the lowest median TTFS alone
Human-read transcripts: captions, notes, search, or analyticslowest WERuniversal-3.6-pro AssemblyAI · 2.1%
Emails, phones, paths, or IDs must be exactnot this boardStructured Speech-to-Text Benchmark

Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. For emails, phones, paths, and IDs, go to the Structured Speech-to-Text Benchmark. Full ranking below.

Run by Coval, a voice-AI evaluation platform

This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution. A public pinned set is a reproducibility target and a contamination risk: continuous re-running against a fixed set does not make that tension smaller over time.

synced from Covalrolling 7d aggregate · pulled 2026-10-09 12:20 UTC · this table bada4aec9614 · pinned dataset 09d98bd6450c (stable; does not move with sample counts) · methodology → · benchmarks.coval.ai → · open-source runner: re-run it yourself →

These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.

Speech-to-Text models ranked by Word Error Rate

Every model runs the same pinned dataset under the same conditions. Last 7d. The Pareto frontier below is the useful cut; the full ranking after it is every model, sorted by average WER.

The Pareto frontier

Same meaning as in the definition at the top: a model is here only if nothing else is both as fast or faster and as accurate or more accurate. Example: whisper-large-v3 and parakeet-tdt-0.6b-v3 are both about 111 ms median TTFS, but whisper-large-v3 is 4.4% WER against 11.0% for parakeet-tdt-0.6b-v3, so parakeet-tdt-0.6b-v3 is not in this table. Currently 3 of 28 models with both metrics. This is not a single winner: the fastest row can have much worse WER.

The Pareto frontier is not a fixed ranking. A week is long enough to change the set. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro. ink-2 and inworld-stt-1 joined between Aug 5 and Aug 10, after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms.

ModelTTFS medianTTFS p95WER avg
qwen3-asr-1.7b
Baseten
21 ms111 ms3.0%
qwen3-asr-fast
Nari
44 ms69 ms2.6%
universal-3.6-pro
AssemblyAI
140 ms286 ms2.1%

Full ranking, sorted by WER

Lower is better. Median TTFS is shown for context.

Top 8 per metric
#ModelTTFS medianTTFS p95WER avgSamples
1universal-3.6-pro
AssemblyAI
140 ms286 ms2.1%6,338
2universal-3.5-pro
AssemblyAI
144 ms351 ms2.2%6,375
3qwen3-asr-fast
Nari
44 ms69 ms2.6%6,375
4realtime
Reson8
237 ms269 ms2.6%6,375
5qwen3-asr-1.7b †
Baseten
21 ms111 ms3.0%3,191
6gemini-3.5-transcribe-live
Gemini
295 ms455 ms3.1%6,360
7chirp_3
Google
753 ms937 ms3.2%6,375
8ink-2
Cartesia
148 ms199 ms3.5%6,375
9inworld-stt-1
Inworld AI
65 ms104 ms3.8%6,362
10chirp_2
Google
793 ms1,442 ms3.9%6,374
11linden-1
Speechmatics
348 ms412 ms3.9%6,375
12gpt-4o-transcribe
OpenAI
750 ms1,047 ms3.9%6,354
13gpt-4o-mini-transcribe
OpenAI
642 ms1,085 ms4.0%6,356
14gpt-realtime-whisper
OpenAI
553 ms698 ms4.2%6,353
15scribe_v2_realtime
ElevenLabs
115 ms183 ms4.3%6,374
16grok-stt
xAI
153 ms274 ms4.4%6,361
17whisper-large-v3 †
Baseten
111 ms245 ms4.4%3,187
18pulse
Smallest
190 ms283 ms4.7%6,375
19stt-rt-v5
Soniox
36 ms66 ms4.9%6,375
20nova-3
Deepgram
90 ms135 ms5.3%6,372
21voxtral-mini-transcribe-realtime-2602
Mistral
330 ms636 ms5.4%6,156
22flux-general-en
Deepgram
99 ms164 ms5.9%6,373
23universal-streaming ‡
AssemblyAI
----6.1%6,367
24solaria-1
Gladia
367 ms736 ms6.3%6,350
25nova-2
Deepgram
92 ms145 ms6.7%6,370
26whisper-large-v3
Together AI
107 ms742 ms7.5%6,342
27flux-general-multi
Deepgram
101 ms165 ms7.7%6,373
28default
Gradium
244 ms325 ms9.0%6,373
29parakeet-tdt-0.6b-v3
Together AI
112 ms221 ms11.0%6,301
30nemotron-3.5-asr-streaming-0.6b ‡
Together AI
----13.3%6,350

† 2 rows have far fewer samples than the rest of the board (typical 6,368.5): qwen3-asr-1.7b (3,191); whisper-large-v3 (3,187). Treat the whole row: WER and TTFS, including p95, as thinner evidence. A p95 on 3,187 samples rests on about 159 tail points. ‡ 2 models have no TTFS in this window. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →

What this snapshot does not tell you

  • Who decides that speech ended

    TTFS is lag after speech ends. If the provider's own VAD makes that call, aggressive endpointing can buy latency by clipping words. The published definition does not rule that out.

  • Text normalization and call region

    How numbers, casing, punctuation, and contractions are handled can move WER by whole points, and is not specified here. At 65–120 ms, which region the runner calls from is a large fraction of the number, and is also unstated.

  • What decides real deployments

    Hallucination on silence, endpointing behavior, timestamps, diarization, long-form drift, cost, and rate limits under load are not on this board. Emails, phones, paths, and IDs are on the Structured Speech-to-Text Benchmark.

Compare two speech-to-text providers directly

Provider vs provider on the same measured data: best model per axis, winner per axis:

Deepgram vs AssemblyAI → · OpenAI vs Deepgram → · Soniox vs Deepgram →

Speech-to-Text benchmark: common questions

Which speech-to-text API is the fastest?

qwen3-asr-1.7b (Baseten) has the lowest median Time to Final Segment in the current 7d window: 21 ms. That is not a pick by itself: its WER is 3.0% against 2.1% for the WER leader. The fastest 8 models on median TTFS span 80 ms. Median gaps of tens of milliseconds are usually smaller than endpointing, network, and the rest of a voice pipeline. For realtime, read p95 TTFS and WER together; start from the Pareto frontier.

Which speech-to-text API is the most accurate?

universal-3.6-pro (AssemblyAI) has the lowest average Word Error Rate: 2.1%. Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.

Is OpenAI (GPT-4o Transcribe) the best STT API?

On the measured axes: OpenAI's fastest model, gpt-realtime-whisper, ranks #24 of 28 on median TTFS (553 ms), and its most accurate, gpt-4o-transcribe, ranks #12 of 30 on WER (3.9%). Lowest median TTFS is qwen3-asr-1.7b (21 ms); lowest WER is universal-3.6-pro (2.1%). Median TTFS is the wrong single ranking for realtime. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.

What is Time to Final Segment (TTFS)?

Time from the end of speech to the final transcript segment being returned. The lag before a downstream task can use what was said. Lower is better. The tail (p95) is what a realtime pipeline feels; median gaps of tens of milliseconds are usually smaller than endpointing, network, and everything after the transcript.

What is Word Error Rate (WER)?

Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.

Does this benchmark measure price?

No. Pricing is not measured. The benchmark measures latency (TTFS) and Word Error Rate on identical audio. Hallucination on silence, endpointing, timestamps, diarization, long-form drift, and rate limits are also unmeasured; check each provider's pricing page for cost.

Is this STT benchmark independent? Who runs it?

Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.

Why is this a live speech-to-text benchmark?

Because the ranking is a rolling 7-day window. Coval re-runs continuously; this page mirrors that window. Models move on latency and WER, and that is enough for new systems to enter the Pareto frontier. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro on the frontier. Between Aug 5 and Aug 10, ink-2 and inworld-stt-1 entered after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms. A static leaderboard would still be showing last week's set. Last pull 2026-10-09 12:20 UTC. The table hash on this page is of the figures shown; a Coval run ID does not identify this window.

Does a low Word Error Rate mean important fields survive for downstream tasks?

No. WER treats every word equally, so a transcript can look clean and still get an email, URL, or ID wrong. This board measures raw Word Error Rate and Time to Final Segment only. If those fields have to be correct for a downstream task, go to the Structured Speech-to-Text Benchmark.

What is the Pareto frontier on this board?

Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.

Why do some models have no TTFS?

2 of 30 models in this window have no Time to Final Segment. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Treat missing latency as undisclosed, not as zero.

Are multilingual and English-only models compared fairly?

They share this board. flux-general-multi sit next to English-only rows. The pinned set is not characterized by language, accent, or domain here. If it is English-dominant, multilingual models can look worse on WER for a reason that is not their English quality. Treat mixed-language rows as a different category.