Independent Live Speech-to-Text (STT/ASR) Benchmark
Why did we build this benchmark?
People pick speech-to-text on two questions: which API is fastest, and which has the lowest Word Error Rate. This board measures those axes on identical audio. They are not one ranking, so we do not crown a single winner. We show the Pareto frontier first: the models that are not beaten on both speed and accuracy at once. Right now that is 3 of 28 models with both metrics; the rest lose on latency and WER together. The fastest 8 models on median TTFS span 80 ms, which is usually smaller than endpointing, network, and the rest of a voice pipeline. It is live because that set can change: models enter the Pareto frontier as the rolling 7-day window moves. A low WER still does not mean important fields such as emails, URLs, or IDs were captured correctly for a downstream task. For that, use the Structured Speech-to-Text Benchmark.
Definition
Headline metrics: TTFS (Time to Final Segment, the lag after speech ends) and WER (share of words wrong). Same pinned dataset, production-realistic conditions, mirrored here daily with attribution. Read p95 TTFS with average WER: the latency tail is what a realtime pipeline feels. Last 7d.
Pareto frontier: Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.
How to pick from this board
Live / realtime transcription
Do not rank by median TTFS. Start from the Pareto frontier, then read p95 TTFS and WER together. Median gaps of tens of milliseconds are usually smaller than the rest of a voice pipeline, and the lowest-median model can have much worse WER.
Human-read transcripts, captions, or search
Rank by Word Error Rate. That is the right axis when a person reads the text.
Emails, phones, paths, or IDs must be exact
Do not pick from this board. Go to the Structured Speech-to-Text Benchmark.
Is there a best STT model?
No, not as a single ranking. Fastest and lowest WER are often different systems, and ranking by median TTFS hides models that are beaten on both axes at once. That is why we show the Pareto frontier. Structured capture (emails, phones, IDs) is a different failure mode on a different board. For a combined buying guide, see the best speech-to-text models by use case.
Which speech-to-text API is best?
There is no single “best”. Lowest median TTFS is qwen3-asr-1.7b (Baseten) at 21 ms with 3.0% WER; lowest WER is universal-3.6-pro (AssemblyAI) at 2.1%. Median TTFS is the wrong single ranking for realtime. Route by the constraint that actually binds:
| Your constraint | Pick by | Current leader |
|---|---|---|
| Live / realtime transcription: the transcript has to land, but tens of milliseconds of median TTFS rarely decide the pipeline | p95 TTFS and WER together | the Pareto frontier (3 models) not the lowest median TTFS alone |
| Human-read transcripts: captions, notes, search, or analytics | lowest WER | universal-3.6-pro AssemblyAI · 2.1% |
| Emails, phones, paths, or IDs must be exact | not this board | Structured Speech-to-Text Benchmark |
Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. For emails, phones, paths, and IDs, go to the Structured Speech-to-Text Benchmark. Full ranking below.
Run by Coval, a voice-AI evaluation platform
This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution. A public pinned set is a reproducibility target and a contamination risk: continuous re-running against a fixed set does not make that tension smaller over time.
These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.
Speech-to-Text models ranked by Word Error Rate
Every model runs the same pinned dataset under the same conditions. Last 7d. The Pareto frontier below is the useful cut; the full ranking after it is every model, sorted by average WER.
The Pareto frontier
Same meaning as in the definition at the top: a model is here only if nothing else is both as fast or faster and as accurate or more accurate. Example: whisper-large-v3 and parakeet-tdt-0.6b-v3 are both about 111 ms median TTFS, but whisper-large-v3 is 4.4% WER against 11.0% for parakeet-tdt-0.6b-v3, so parakeet-tdt-0.6b-v3 is not in this table. Currently 3 of 28 models with both metrics. This is not a single winner: the fastest row can have much worse WER.
The Pareto frontier is not a fixed ranking. A week is long enough to change the set. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro. ink-2 and inworld-stt-1 joined between Aug 5 and Aug 10, after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms.
| Model | TTFS median | TTFS p95 | WER avg |
|---|---|---|---|
| qwen3-asr-1.7b Baseten | 21 ms | 111 ms | 3.0% |
| qwen3-asr-fast Nari | 44 ms | 69 ms | 2.6% |
| universal-3.6-pro AssemblyAI | 140 ms | 286 ms | 2.1% |
Full ranking, sorted by WER
Lower is better. Median TTFS is shown for context.
| # | Model | TTFS median | TTFS p95 | WER avg | Samples |
|---|---|---|---|---|---|
| 1 | universal-3.6-pro AssemblyAI | 140 ms | 286 ms | 2.1% | 6,338 |
| 2 | universal-3.5-pro AssemblyAI | 144 ms | 351 ms | 2.2% | 6,375 |
| 3 | qwen3-asr-fast Nari | 44 ms | 69 ms | 2.6% | 6,375 |
| 4 | realtime Reson8 | 237 ms | 269 ms | 2.6% | 6,375 |
| 5 | qwen3-asr-1.7b † Baseten | 21 ms | 111 ms | 3.0% | 3,191 |
| 6 | gemini-3.5-transcribe-live Gemini | 295 ms | 455 ms | 3.1% | 6,360 |
| 7 | chirp_3 | 753 ms | 937 ms | 3.2% | 6,375 |
| 8 | ink-2 Cartesia | 148 ms | 199 ms | 3.5% | 6,375 |
| 9 | inworld-stt-1 Inworld AI | 65 ms | 104 ms | 3.8% | 6,362 |
| 10 | chirp_2 | 793 ms | 1,442 ms | 3.9% | 6,374 |
| 11 | linden-1 Speechmatics | 348 ms | 412 ms | 3.9% | 6,375 |
| 12 | gpt-4o-transcribe OpenAI | 750 ms | 1,047 ms | 3.9% | 6,354 |
| 13 | gpt-4o-mini-transcribe OpenAI | 642 ms | 1,085 ms | 4.0% | 6,356 |
| 14 | gpt-realtime-whisper OpenAI | 553 ms | 698 ms | 4.2% | 6,353 |
| 15 | scribe_v2_realtime ElevenLabs | 115 ms | 183 ms | 4.3% | 6,374 |
| 16 | grok-stt xAI | 153 ms | 274 ms | 4.4% | 6,361 |
| 17 | whisper-large-v3 † Baseten | 111 ms | 245 ms | 4.4% | 3,187 |
| 18 | pulse Smallest | 190 ms | 283 ms | 4.7% | 6,375 |
| 19 | stt-rt-v5 Soniox | 36 ms | 66 ms | 4.9% | 6,375 |
| 20 | nova-3 Deepgram | 90 ms | 135 ms | 5.3% | 6,372 |
| 21 | voxtral-mini-transcribe-realtime-2602 Mistral | 330 ms | 636 ms | 5.4% | 6,156 |
| 22 | flux-general-en Deepgram | 99 ms | 164 ms | 5.9% | 6,373 |
| 23 | universal-streaming ‡ AssemblyAI | -- | -- | 6.1% | 6,367 |
| 24 | solaria-1 Gladia | 367 ms | 736 ms | 6.3% | 6,350 |
| 25 | nova-2 Deepgram | 92 ms | 145 ms | 6.7% | 6,370 |
| 26 | whisper-large-v3 Together AI | 107 ms | 742 ms | 7.5% | 6,342 |
| 27 | flux-general-multi Deepgram | 101 ms | 165 ms | 7.7% | 6,373 |
| 28 | default Gradium | 244 ms | 325 ms | 9.0% | 6,373 |
| 29 | parakeet-tdt-0.6b-v3 Together AI | 112 ms | 221 ms | 11.0% | 6,301 |
| 30 | nemotron-3.5-asr-streaming-0.6b ‡ Together AI | -- | -- | 13.3% | 6,350 |
† 2 rows have far fewer samples than the rest of the board (typical 6,368.5): qwen3-asr-1.7b (3,191); whisper-large-v3 (3,187). Treat the whole row: WER and TTFS, including p95, as thinner evidence. A p95 on 3,187 samples rests on about 159 tail points. ‡ 2 models have no TTFS in this window. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →
What this snapshot does not tell you
Who decides that speech ended
TTFS is lag after speech ends. If the provider's own VAD makes that call, aggressive endpointing can buy latency by clipping words. The published definition does not rule that out.
Text normalization and call region
How numbers, casing, punctuation, and contractions are handled can move WER by whole points, and is not specified here. At 65–120 ms, which region the runner calls from is a large fraction of the number, and is also unstated.
What decides real deployments
Hallucination on silence, endpointing behavior, timestamps, diarization, long-form drift, cost, and rate limits under load are not on this board. Emails, phones, paths, and IDs are on the Structured Speech-to-Text Benchmark.
Compare two speech-to-text providers directly
Provider vs provider on the same measured data: best model per axis, winner per axis:
Deepgram vs AssemblyAI → · OpenAI vs Deepgram → · Soniox vs Deepgram →
Speech-to-Text benchmark: common questions
Which speech-to-text API is the fastest?
qwen3-asr-1.7b (Baseten) has the lowest median Time to Final Segment in the current 7d window: 21 ms. That is not a pick by itself: its WER is 3.0% against 2.1% for the WER leader. The fastest 8 models on median TTFS span 80 ms. Median gaps of tens of milliseconds are usually smaller than endpointing, network, and the rest of a voice pipeline. For realtime, read p95 TTFS and WER together; start from the Pareto frontier.
Which speech-to-text API is the most accurate?
universal-3.6-pro (AssemblyAI) has the lowest average Word Error Rate: 2.1%. Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.
Is OpenAI (GPT-4o Transcribe) the best STT API?
On the measured axes: OpenAI's fastest model, gpt-realtime-whisper, ranks #24 of 28 on median TTFS (553 ms), and its most accurate, gpt-4o-transcribe, ranks #12 of 30 on WER (3.9%). Lowest median TTFS is qwen3-asr-1.7b (21 ms); lowest WER is universal-3.6-pro (2.1%). Median TTFS is the wrong single ranking for realtime. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.
What is Time to Final Segment (TTFS)?
Time from the end of speech to the final transcript segment being returned. The lag before a downstream task can use what was said. Lower is better. The tail (p95) is what a realtime pipeline feels; median gaps of tens of milliseconds are usually smaller than endpointing, network, and everything after the transcript.
What is Word Error Rate (WER)?
Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.
Does this benchmark measure price?
No. Pricing is not measured. The benchmark measures latency (TTFS) and Word Error Rate on identical audio. Hallucination on silence, endpointing, timestamps, diarization, long-form drift, and rate limits are also unmeasured; check each provider's pricing page for cost.
Is this STT benchmark independent? Who runs it?
Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.
Why is this a live speech-to-text benchmark?
Because the ranking is a rolling 7-day window. Coval re-runs continuously; this page mirrors that window. Models move on latency and WER, and that is enough for new systems to enter the Pareto frontier. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro on the frontier. Between Aug 5 and Aug 10, ink-2 and inworld-stt-1 entered after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms. A static leaderboard would still be showing last week's set. Last pull 2026-10-09 12:20 UTC. The table hash on this page is of the figures shown; a Coval run ID does not identify this window.
Does a low Word Error Rate mean important fields survive for downstream tasks?
No. WER treats every word equally, so a transcript can look clean and still get an email, URL, or ID wrong. This board measures raw Word Error Rate and Time to Final Segment only. If those fields have to be correct for a downstream task, go to the Structured Speech-to-Text Benchmark.
What is the Pareto frontier on this board?
Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.
Why do some models have no TTFS?
2 of 30 models in this window have no Time to Final Segment. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Treat missing latency as undisclosed, not as zero.
Are multilingual and English-only models compared fairly?
They share this board. flux-general-multi sit next to English-only rows. The pinned set is not characterized by language, accent, or domain here. If it is English-dominant, multilingual models can look worse on WER for a reason that is not their English quality. Treat mixed-language rows as a different category.