Independent Live Speech-to-Text (STT/ASR) Benchmark
Why did we build this benchmark?
People pick speech-to-text on two questions: which API is fastest, and which has the lowest Word Error Rate. This board measures those axes on identical audio. They are not one ranking, so we do not crown a single winner. We show the Pareto frontier first: the models that are not beaten on both speed and accuracy at once. Right now that is 4 of 24 models with both metrics; the rest lose on latency and WER together. The fastest 11 models on median TTFS span 79 ms, which is usually smaller than endpointing, network, and the rest of a voice pipeline. It is live because that set can change: models enter the Pareto frontier as the rolling 7-day window moves. A low WER still does not mean important fields such as emails, URLs, or IDs were captured correctly for a downstream task. For that, use the Structured Speech-to-Text Benchmark.
Definition
Headline metrics: TTFS (Time to Final Segment, the lag after speech ends) and WER (share of words wrong). Same pinned dataset, production-realistic conditions, mirrored here daily with attribution. Read p95 TTFS with average WER: the latency tail is what a realtime pipeline feels. Last 7d.
Pareto frontier: Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.
How to pick from this board
Live / realtime transcription
Do not rank by median TTFS. Start from the Pareto frontier, then read p95 TTFS and WER together. Median gaps of tens of milliseconds are usually smaller than the rest of a voice pipeline, and the lowest-median model can have much worse WER.
Human-read transcripts, captions, or search
Rank by Word Error Rate. That is the right axis when a person reads the text.
Emails, phones, paths, or IDs must be exact
Do not pick from this board. Go to the Structured Speech-to-Text Benchmark.
Is there a best STT model?
No, not as a single ranking. Fastest and lowest WER are often different systems, and ranking by median TTFS hides models that are beaten on both axes at once. That is why we show the Pareto frontier. Structured capture (emails, phones, IDs) is a different failure mode on a different board. For a combined buying guide, see the best speech-to-text models by use case.
Which speech-to-text API is best?
There is no single “best”. Lowest median TTFS is stt-rt-v5 (Soniox) at 49 ms with 5.7% WER; lowest WER is universal-3.5-pro (AssemblyAI) at 3.3%. Median TTFS is the wrong single ranking for realtime. Route by the constraint that actually binds:
| Your constraint | Pick by | Current leader |
|---|---|---|
| Live / realtime transcription: the transcript has to land, but tens of milliseconds of median TTFS rarely decide the pipeline | p95 TTFS and WER together | the Pareto frontier (4 models) not the lowest median TTFS alone |
| Human-read transcripts: captions, notes, search, or analytics | lowest WER | universal-3.5-pro AssemblyAI · 3.3% |
| Emails, phones, paths, or IDs must be exact | not this board | Structured Speech-to-Text Benchmark |
Not measured here: price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim. For emails, phones, paths, and IDs, go to the Structured Speech-to-Text Benchmark. Full ranking below.
Run by Coval, a voice-AI evaluation platform
This benchmark is produced by Coval, which sells voice-agent evaluation infrastructure. Every number here is Coval's, measured on a pinned dataset under production-realistic conditions; the methodology, runner code, and inputs are open-source (Apache-2.0) and reproducible by re-running the suite. Openbenchmarks mirrors these results with attribution. A public pinned set is a reproducibility target and a contamination risk: continuous re-running against a fixed set does not make that tension smaller over time.
These figures are a rolling 7d window, not one job. A Coval run ID is not printed here because it does not identify this table. The table hash is SHA-256 of the numbers on this page. The dataset hash is the pinned audio set.
Speech-to-Text models ranked by Word Error Rate
Every model runs the same pinned dataset under the same conditions. Last 7d. The Pareto frontier below is the useful cut; the full ranking after it is every model, sorted by average WER.
The Pareto frontier
Same meaning as in the definition at the top: a model is here only if nothing else is both as fast or faster and as accurate or more accurate. Example: nova-3 and nova-2 are both about 95 ms median TTFS, but nova-3 is 6.3% WER against 8.0% for nova-2, so nova-2 is not in this table. Currently 4 of 24 models with both metrics. This is not a single winner: the fastest row can have much worse WER.
The Pareto frontier is not a fixed ranking. A week is long enough to change the set. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro. ink-2 and inworld-stt-1 joined between Aug 5 and Aug 10, after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms.
| Model | TTFS median | TTFS p95 | WER avg |
|---|---|---|---|
| stt-rt-v5 Soniox | 49 ms | 77 ms | 5.7% |
| inworld-stt-1 Inworld AI | 63 ms | 89 ms | 5.1% |
| ink-2 Cartesia | 105 ms | 138 ms | 4.8% |
| universal-3.5-pro AssemblyAI | 128 ms | 317 ms | 3.3% |
Full ranking, sorted by WER
Lower is better. Median TTFS is shown for context.
| # | Model | TTFS median | TTFS p95 | WER avg | Samples |
|---|---|---|---|---|---|
| 1 | universal-3.5-pro AssemblyAI | 128 ms | 317 ms | 3.3% | 6,787 |
| 2 | realtime Reson8 | 266 ms | 300 ms | 3.5% | 6,782 |
| 3 | chirp_3 | 780 ms | 1,001 ms | 4.2% | 6,788 |
| 4 | enhanced Speechmatics | 324 ms | 512 ms | 4.3% | 6,785 |
| 5 | gpt-4o-transcribe OpenAI | 704 ms | 954 ms | 4.6% | 6,786 |
| 6 | grok-stt xAI | 194 ms | 263 ms | 4.7% | 6,787 |
| 7 | ink-2 Cartesia | 105 ms | 138 ms | 4.8% | 6,781 |
| 8 | velma-2-stt-streaming-english-v2 † Modulate | 123 ms | 177 ms | 4.9% | 2,618 |
| 9 | chirp_2 | 784 ms | 1,286 ms | 4.9% | 6,752 |
| 10 | gpt-4o-mini-transcribe OpenAI | 621 ms | 942 ms | 5.0% | 6,786 |
| 11 | inworld-stt-1 Inworld AI | 63 ms | 89 ms | 5.1% | 6,775 |
| 12 | gpt-realtime-whisper OpenAI | 557 ms | 673 ms | 5.2% | 6,777 |
| 13 | default Speechmatics | 217 ms | 273 ms | 5.4% | 6,787 |
| 14 | pulse Smallest | 194 ms | 292 ms | 5.4% | 6,786 |
| 15 | scribe_v2_realtime ElevenLabs | 111 ms | 195 ms | 5.6% | 6,785 |
| 16 | stt-rt-v5 Soniox | 49 ms | 77 ms | 5.7% | 6,787 |
| 17 | voxtral-mini-transcribe-realtime-2602 Mistral | 325 ms | 536 ms | 6.1% | 6,762 |
| 18 | nova-3 Deepgram | 95 ms | 141 ms | 6.3% | 6,780 |
| 19 | velma-2-stt-streaming † Modulate | 92 ms | 1,028 ms | 6.5% | 2,604 |
| 20 | flux-general-en ‡ Deepgram | -- | -- | 6.5% | 6,787 |
| 21 | universal-streaming ‡ AssemblyAI | -- | -- | 7.0% | 6,771 |
| 22 | solaria-1 Gladia | 398 ms | 2,574 ms | 7.3% | 6,767 |
| 23 | nova-2 Deepgram | 96 ms | 155 ms | 8.0% | 6,777 |
| 24 | whisper-large-v3 Together AI | 122 ms | 417 ms | 8.2% | 6,770 |
| 25 | flux-general-multi ‡ Deepgram | -- | -- | 8.5% | 6,782 |
| 26 | default Gradium | 248 ms | 326 ms | 10.1% | 6,780 |
| 27 | parakeet-tdt-0.6b-v3 Together AI | 62 ms | 137 ms | 11.2% | 6,764 |
| 28 | universal-streaming-multilingual † ‡ AssemblyAI | -- | -- | 11.8% | 2,589 |
| 29 | nemotron-3.5-asr-streaming-0.6b ‡ Together AI | -- | -- | 15.5% | 6,737 |
† 3 rows have far fewer samples than the rest of the board (typical 6,780): velma-2-stt-streaming (2,604); velma-2-stt-streaming-english-v2 (2,618); universal-streaming-multilingual (2,589). Treat the whole row: WER and TTFS, including p95, as thinner evidence. A p95 on 2,589 samples rests on about 129 tail points. ‡ 5 models have no TTFS in this window. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Numbers are point-in-time against Coval's pinned dataset and refresh continuously. Full charts, distributions, and windows (24h / 7d / 30d) at benchmarks.coval.ai →
What this snapshot does not tell you
Who decides that speech ended
TTFS is lag after speech ends. If the provider's own VAD makes that call, aggressive endpointing can buy latency by clipping words. The published definition does not rule that out.
Text normalization and call region
How numbers, casing, punctuation, and contractions are handled can move WER by whole points, and is not specified here. At 65–120 ms, which region the runner calls from is a large fraction of the number, and is also unstated.
What decides real deployments
Hallucination on silence, endpointing behavior, timestamps, diarization, long-form drift, cost, and rate limits under load are not on this board. Emails, phones, paths, and IDs are on the Structured Speech-to-Text Benchmark.
Compare two speech-to-text providers directly
Provider vs provider on the same measured data: best model per axis, winner per axis:
Deepgram vs AssemblyAI → · OpenAI vs Deepgram → · Soniox vs Deepgram →
Speech-to-Text benchmark: common questions
Which speech-to-text API is the fastest?
stt-rt-v5 (Soniox) has the lowest median Time to Final Segment in the current 7d window: 49 ms. That is not a pick by itself: its WER is 5.7% against 3.3% for the WER leader. The fastest 11 models on median TTFS span 79 ms. Median gaps of tens of milliseconds are usually smaller than endpointing, network, and the rest of a voice pipeline. For realtime, read p95 TTFS and WER together; start from the Pareto frontier.
Which speech-to-text API is the most accurate?
universal-3.5-pro (AssemblyAI) has the lowest average Word Error Rate: 3.3%. Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.
Is OpenAI (GPT-4o Transcribe) the best STT API?
On the measured axes: OpenAI's fastest model, gpt-realtime-whisper, ranks #20 of 24 on median TTFS (557 ms), and its most accurate, gpt-4o-transcribe, ranks #5 of 29 on WER (4.6%). Lowest median TTFS is stt-rt-v5 (49 ms); lowest WER is universal-3.5-pro (3.3%). Median TTFS is the wrong single ranking for realtime. price, hallucination on silence, endpointing, timestamps, diarization, long-form drift, rate limits, and whether emails, phones, paths, and IDs survive verbatim are not measured here.
What is Time to Final Segment (TTFS)?
Time from the end of speech to the final transcript segment being returned. The lag before a downstream task can use what was said. Lower is better. The tail (p95) is what a realtime pipeline feels; median gaps of tens of milliseconds are usually smaller than endpointing, network, and everything after the transcript.
What is Word Error Rate (WER)?
Share of words wrong in the transcript (substitutions + insertions + deletions, divided by words spoken). Lower is better. WER treats every word equally, so it does not tell you whether an email, phone number, or ID survived.
Does this benchmark measure price?
No. Pricing is not measured. The benchmark measures latency (TTFS) and Word Error Rate on identical audio. Hallucination on silence, endpointing, timestamps, diarization, long-form drift, and rate limits are also unmeasured; check each provider's pricing page for cost.
Is this STT benchmark independent? Who runs it?
Coval sells voice-agent evaluation infrastructure. The methodology, runner code, and pinned dataset are open-source (Apache-2.0) and re-runnable end to end; Openbenchmarks mirrors the results with attribution.
Why is this a live speech-to-text benchmark?
Because the ranking is a rolling 7-day window. Coval re-runs continuously; this page mirrors that window. Models move on latency and WER, and that is enough for new systems to enter the Pareto frontier. Aug 1 and Aug 5 had only parakeet-tdt-0.6b-v3, stt-rt-v5, and universal-3.5-pro on the frontier. Between Aug 5 and Aug 10, ink-2 and inworld-stt-1 entered after universal-3.5-pro's median TTFS moved from about 99 ms to 111 ms. A static leaderboard would still be showing last week's set. Last pull 2026-08-25 06:31 UTC. The table hash on this page is of the figures shown; a Coval run ID does not identify this window.
Does a low Word Error Rate mean important fields survive for downstream tasks?
No. WER treats every word equally, so a transcript can look clean and still get an email, URL, or ID wrong. This board measures raw Word Error Rate and Time to Final Segment only. If those fields have to be correct for a downstream task, go to the Structured Speech-to-Text Benchmark.
What is the Pareto frontier on this board?
Models that are not beaten on both speed and accuracy at once. A model is on it only if nothing else is both as fast or faster (median TTFS) and as accurate or more accurate (WER). Improving one axis from a frontier model means giving up the other. We show it instead of a single ranking because ranking by median TTFS or by WER alone makes the order look like a race. Most models lose on speed and accuracy together; those rows still appear in the full ranking, but they are not on the frontier. The cut is a rolling 7d window: membership can hold for days and still change over a week.
Why do some models have no TTFS?
5 of 29 models in this window have no Time to Final Segment. The snapshot does not say whether those APIs are batch-only, do not signal segment finality, or failed measurement. Treat missing latency as undisclosed, not as zero.
Are multilingual and English-only models compared fairly?
They share this board. universal-streaming-multilingual, flux-general-multi sit next to English-only rows. The pinned set is not characterized by language, accent, or domain here. If it is English-dominant, multilingual models can look worse on WER for a reason that is not their English quality. Treat mixed-language rows as a different category.