Independent Structured Speech-to-Text (STT/ASR) Benchmark
Why did we build this benchmark?
As a standard metric, Word Error Rate (WER) treats every word as equal. A transcript can sound clean and still corrupt the callback number, email or URL that a downstream task needs verbatim. We built this benchmark on a dataset provided by Besimple AI to measure that failure mode and to show that WER and task success can decouple. On this snapshot AssemblyAI Universal 3.5 Pro leads WER at 3.9% but ranks 11 of 17 on Task Success Rate (50.3%), while Deepgram Nova-3 and ElevenLabs Scribe v2 are statistically tied at the top on Task Success Rate (71.0% vs 68.0%, within roughly ±2.6pp sampling noise on 300 recordings).
Definition & problem
This is an ASR/STT benchmark of 17 systems, including Deepgram, ElevenLabs, Google Cloud, OpenAI, and Thinking Machines Lab, on 300 human-recorded segments (5.587 hours, 85 speakers) with 1,482 audited target entities across 26 types. Every system gets the same raw audio and nothing else: no prompts, custom vocabulary, or post-ASR cleanup. A versioned LLM verifier (gpt-5.6-sol) judges each entity. Headline metrics: TSR (recordings with every value correct) and CTEM (values recovered exactly), with format-invariant WER as the broad diagnostic.
How to pick the best STT for your workflow
Booking appointments or support callbacks
The task must capture a phone number, email, date, and time exactly, since one wrong digit fails the run. TSR is a fair first screen for that job; confirm on the phone, email, date, and time rows in the entity matrix before you ship.
Voice coding tasks
Do not pick from headline TSR. Rank the command, CLI-flag, file-path, and environment-variable rows in the entity matrix (those leaders are often not the TSR leaders).
CRM or form filling from a call
Account numbers, reference IDs, and postal addresses are the failure modes that matter, so read those entity rows rather than only the headline rank.
Live / realtime tasks vs post-call batch
If realtime looks fine in demos but production streams drop digits, compare the same vendor's batch and streaming rows.
Is there a best STT model?
No, not as a single ranking. “Best STT” only makes sense once you name the use case and the failure mode you care about.
Downstream task or software that needs every field right
Booking, callbacks, form fill: use Task Success Rate on this board. The top systems are often statistically tied within sampling noise; confirm on the entity rows for phones, emails, dates, and IDs.
Human-read transcripts, cheapest, or fastest
That is a raw Word Error Rate and latency question. Use the Speech-to-Text benchmark, not this page.
ChangelogLast updated Aug 7, 2026 · 2026 Q3 snapshot: 17 systems (batch + streaming), including Inworld STT 1 and self-hosted Parakeet / OmniASR baselines; all re-verified with gpt-5.6-sol.
Leaderboard and per-entity-type breakdown
Measured TSR, CTEM, and format-invariant WER on identical raw audio. Batch and streaming endpoints are separate rows. No opinions in this section: numbers only.
| Rank | System | Vendor | Mode | Deployment | TSR | CTEM | WER |
|---|---|---|---|---|---|---|---|
| 1 | Deepgram Nova-3 | Deepgram | batch | hosted API | 71.0% | 91.4% | 8.2% |
| 2 | ElevenLabs Scribe v2 | ElevenLabs | batch | hosted API | 68.0% | 91.8% | 6.8% |
| 3 | Deepgram Nova-3 | Deepgram | streaming | hosted API | 61.0% | 88.3% | 9.5% |
| 4 | ElevenLabs Scribe v2 Realtime | ElevenLabs | streaming | hosted API | 59.0% | 88.1% | 5.1% |
| 5 | Google Cloud Chirp 3 | Google Cloud | batch | hosted API | 58.0% | 88.7% | 5.8% |
| 6 | OpenAI GPT Realtime (Whisper) | OpenAI | streaming | hosted API | 57.3% | 88.7% | 8.6% |
| 7 | Google Cloud Chirp 3 | Google Cloud | streaming | hosted API | 56.3% | 87.3% | 6.8% |
| 8 | Inkling (Thinking Machines) | Thinking Machines Lab | batch | managed endpoint via Modal | 54.7% | 85.3% | 5.0% |
| 9 | OpenAI GPT-4o Transcribe | OpenAI | batch | hosted API | 54.0% | 87.0% | 4.6% |
| 10 | Whisper large-v3 | OpenAI | batch | hosted API | 53.3% | 87.3% | 5.9% |
| 11 | AssemblyAI Universal 3.5 Pro | AssemblyAI | batch | hosted API | 50.3% | 86.2% | 3.9% |
| 12 | Inworld STT 1 | Inworld | batch | hosted API | 48.3% | 85.4% | 8.2% |
| 13 | NVIDIA Parakeet TDT 0.6B v3 | NVIDIA | batch | self-hosted via Modal | 39.3% | 79.8% | 11.4% |
| 14 | Amazon Transcribe | Amazon Web Services | streaming | hosted API | 35.0% | 78.8% | 7.5% |
| 15 | Meta OmniASR LLM Unlimited 7B v2 | Meta | batch | self-hosted via Modal | 32.7% | 76.9% | 10.5% |
| 16 | AssemblyAI Universal 3 Pro | AssemblyAI | streaming | hosted API | 24.7% | 71.9% | 7.5% |
| 17 | Inworld STT 1 | Inworld | streaming | hosted API | 7.3% | 34.0% | 58.2% |
| System | Email Address | Phone Number | Phone Extension | Person Or Team Name | Postal Address | URL | IP Address | Port Number | Command | CLI Flag | File Path | Environment Variable | Code Symbol | Version | Reference ID | Product Code | Account Or Record Number | Currency Amount | Percentage | Measurement | Plain Number | Date | Time | Acronym Or Initialism | Spelled Sequence | Domain Term |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Deepgram Nova-3 · batch | 77% | 98% | 100% | 100% | 70% | 69% | 96% | 97% | 56% | 91% | 61% | 86% | 86% | 93% | 95% | 96% | 97% | 100% | 100% | 100% | 100% | 100% | 97% | 99% | 98% | 90% |
| ElevenLabs Scribe v2 · batch | 69% | 100% | 100% | 100% | 73% | 65% | 96% | 100% | 72% | 93% | 65% | 97% | 97% | 93% | 93% | 91% | 94% | 100% | 100% | 100% | 100% | 99% | 98% | 96% | 98% | 100% |
| Deepgram Nova-3 · streaming | 69% | 98% | 100% | 100% | 68% | 63% | 92% | 90% | 60% | 89% | 53% | 80% | 77% | 93% | 93% | 90% | 92% | 96% | 100% | 98% | 98% | 96% | 97% | 95% | 96% | 90% |
| ElevenLabs Scribe v2 Realtime · streaming | 71% | 98% | 100% | 100% | 60% | 47% | 96% | 97% | 50% | 84% | 47% | 54% | 97% | 100% | 95% | 92% | 92% | 92% | 100% | 100% | 100% | 100% | 98% | 99% | 98% | 100% |
| Google Cloud Chirp 3 · batch | 66% | 98% | 100% | 100% | 57% | 61% | 96% | 97% | 54% | 91% | 59% | 89% | 91% | 90% | 91% | 92% | 94% | 97% | 98% | 100% | 100% | 100% | 93% | 96% | 92% | 90% |
| OpenAI GPT Realtime (Whisper) · streaming | 65% | 90% | 93% | 94% | 53% | 58% | 84% | 97% | 74% | 95% | 67% | 91% | 94% | 90% | 93% | 93% | 91% | 89% | 100% | 98% | 98% | 100% | 98% | 98% | 91% | 95% |
| Google Cloud Chirp 3 · streaming | 68% | 97% | 100% | 100% | 55% | 55% | 96% | 97% | 54% | 86% | 63% | 83% | 83% | 87% | 89% | 96% | 88% | 93% | 98% | 98% | 98% | 100% | 92% | 94% | 94% | 95% |
| Inkling (Thinking Machines) · batch | 57% | 92% | 100% | 89% | 57% | 55% | 84% | 83% | 58% | 80% | 59% | 89% | 83% | 97% | 88% | 83% | 86% | 95% | 98% | 100% | 97% | 100% | 95% | 92% | 93% | 95% |
| OpenAI GPT-4o Transcribe · batch | 63% | 100% | 100% | 94% | 53% | 45% | 80% | 97% | 56% | 86% | 57% | 94% | 89% | 93% | 90% | 91% | 88% | 100% | 100% | 100% | 98% | 100% | 98% | 94% | 91% | 85% |
| Whisper large-v3 · batch | 78% | 95% | 100% | 94% | 60% | 42% | 92% | 90% | 46% | 89% | 53% | 89% | 91% | 93% | 93% | 89% | 91% | 97% | 100% | 100% | 95% | 97% | 90% | 96% | 97% | 85% |
| AssemblyAI Universal 3.5 Pro · batch | 55% | 90% | 97% | 100% | 68% | 47% | 88% | 97% | 46% | 86% | 63% | 83% | 100% | 83% | 91% | 86% | 88% | 92% | 98% | 100% | 97% | 100% | 98% | 93% | 95% | 100% |
| Inworld STT 1 · batch | 60% | 90% | 97% | 89% | 55% | 47% | 72% | 90% | 58% | 84% | 67% | 89% | 91% | 90% | 86% | 89% | 91% | 96% | 96% | 95% | 98% | 98% | 93% | 99% | 89% | 90% |
| NVIDIA Parakeet TDT 0.6B v3 · batch | 29% | 83% | 100% | 94% | 45% | 47% | 76% | 83% | 38% | 86% | 37% | 83% | 83% | 93% | 81% | 81% | 82% | 85% | 100% | 100% | 98% | 97% | 93% | 94% | 90% | 85% |
| Amazon Transcribe · streaming | 28% | 98% | 90% | 83% | 53% | 40% | 88% | 87% | 12% | 61% | 37% | 57% | 77% | 87% | 89% | 86% | 86% | 85% | 98% | 90% | 98% | 99% | 98% | 95% | 87% | 95% |
| Meta OmniASR LLM Unlimited 7B v2 · batch | 35% | 73% | 100% | 89% | 48% | 50% | 72% | 87% | 32% | 70% | 24% | 77% | 80% | 73% | 88% | 86% | 85% | 75% | 50% | 90% | 98% | 94% | 92% | 99% | 96% | 80% |
| AssemblyAI Universal 3 Pro · streaming | 38% | 78% | 90% | 94% | 50% | 21% | 60% | 97% | 34% | 73% | 31% | 69% | 77% | 80% | 63% | 72% | 63% | 73% | 96% | 97% | 95% | 90% | 92% | 91% | 79% | 95% |
| Inworld STT 1 · streaming | 9% | 18% | 17% | 72% | 30% | 27% | 32% | 33% | 18% | 23% | 35% | 29% | 31% | 57% | 55% | 31% | 60% | 12% | 6% | 41% | 40% | 46% | 25% | 29% | 53% | 10% |
[02] methodology and what is not measured+
What each system is asked to do
- The dataset is 300 human-recorded English WAV segments (5.587 hours, 85 anonymized speakers) across 8 workplace workflow domains: support callbacks, deployments, order handling, and similar.
- Each item carries three transcript layers:
template(script with entity placeholders),acoustic(what the speaker says aloud), andcanonical(the written value a downstream application needs). “double dash dry dash run” maps to--dry-run; “all caps database underscore URL” maps toDATABASE_URL. - 1,482 audited target entities span 26 structured types: emails, phone numbers, URLs, IP addresses, ports, shell commands, CLI flags, file paths, environment variables, code symbols, versions, reference IDs, amounts, dates, times, spelled sequences, and more.
- The system's only input is the audio file. Its transcript is stored verbatim with run metadata (endpoint, evaluation date, inference settings).
What each metric means
CTEM· Canonical Token/Entity Match. Correct target entities ÷ target entities. An entity is correct only when the exact canonical value is recoverable; formatting differences are fine, wrong or missing characters are not.TSR· Task Success Rate. Recordings with all target entities correct ÷ recordings. The strictest headline metric: one corrupted digit fails the whole recording. With 300 clips, standard error on TSR is roughly ±2.6pp, so gaps of a few points at the top are noise, not a stable ranking.WER· format-invariant Word Error Rate. Every target entity span accepts either the documented acoustic rendering or the canonical rendering as reference, so212-555-0100and “two one two five five five zero one zero zero” score identically. Non-entity words use ordinary word-level edit distance.per-entity CTEM· CTEM restricted to one of the 26 entity types: the matrix above.
Raw audio only with no hints or cleanup
Every system is evaluated under a raw-audio-only protocol: it receives the audio file and nothing else. Benchmark-specific prompts, target entity lists, domain labels, custom vocabulary, grammar constraints, candidate values, and post-ASR correction are all excluded from the main setting. This measures what the model does on its own, not what a tuned integration could extract from it.
Batch systems receive each recording as a complete file at the provider's documented input format. Streaming systems receive headerless PCM paced in realtime (100 ms chunks) over the provider's websocket protocol, with VAD, turn detection, and vendor canonicalization disabled where the API allows. Batch and streaming endpoints of the same vendor are tracked as separate systems because their error profiles differ substantially.
How entity matches are judged
For each recording, every target entity is checked against the system transcript by a versioned LLM verifier: openai_gpt_5_6_sol_group10_sync_v1 (gpt-5.6-sol, 10 datapoints per request, reasoning effort medium). The verifier marks an entity present only when enough evidence exists to recover the exact canonical value. The same verifier scores every system; its prompt, schema, and settings are frozen into a config digest that ships with every entity-match artifact. A single-judge setup has its own error floor; that is a known limitation, not hidden.
Out of scope on this board
- Multilingual / non-English accuracy (English-only snapshot).
- Diarization, speaker ID, timestamp accuracy, punctuation / formatting quality.
- Noise robustness at specific SNRs, telephony 8 kHz, far-field microphones.
- Domain-specific vocabularies (medical, legal, aviation, heavy proper nouns) beyond the released workplace domains.
- Effect of custom vocabulary / keyword boosting is excluded from the main setting by design.
- On-device / mobile deployment, hallucination on silence, HIPAA / residency, long-form behavior beyond the released clips.
- Raw Word Error Rate and STT latency (time to final segment): not the job of this board (see the sibling Speech-to-Text benchmark for those shortlists).
Verify any number end-to-end
Everything behind this table is released in openbenchmarks-labs/voice-code-bench: audio, transcripts and entity annotations, per-system transcripts plus entity-match decisions, verifier configs, and the scoring CLI. The dataset is also on Hugging Face as besimple-ai/voice-code-bench (MIT). The release script rescores published artifacts and rewrites baselines/results.csv, the table on this page, without provider credentials.
Structured Speech-to-Text Benchmark: common questions
Scope, independence, and how this board differs from raw WER leaderboards. For workflow picks (booking, coding, form fill), use the sections above the tables.
What is the Structured Speech-to-Text Benchmark?
The Structured Speech-to-Text Benchmark is an independent speech-to-text (ASR) benchmark that measures whether a transcript preserves exact structured values in English workplace speech: emails, phone numbers, CLI flags, file paths, URLs, IDs, and 20 other entity types. It scores 17 systems on 300 human-recorded segments (1,482 audited entities) using Task Success Rate (TSR), Canonical Token/Entity Match (CTEM), and format-invariant WER. Dataset by Besimple AI; baselines run and scored by Openbenchmarks.
Does a low Word Error Rate mean the transcript is good enough for software?
No. WER treats every word equally, so a system can post excellent WER while corrupting the one email, URL, or CLI flag a downstream system needs. On this board AssemblyAI Universal 3.5 Pro has the best WER (3.9%) but recovers every structured value in only 50.3% of recordings, while Deepgram Nova-3 and ElevenLabs Scribe v2 sit at the top of TSR (71.0% vs 68.0%). That decoupling is why this board exists.
Where should I go for raw WER or the fastest STT model?
For raw transcript quality (Word Error Rate) and speech-to-text latency, including unqualified “best STT model” shortlists, use the Speech-to-Text benchmark. Use this Structured Speech-to-Text Benchmark when you need the best STT for downstream tasks or software that depends on exact structured values.
Who created the dataset, and is the benchmark independent?
The recordings and entity annotations are authored and released by Besimple AI as besimple-ai/voice-code-bench on Hugging Face (MIT). Openbenchmarks independently runs the 17 tracked systems, scores them with a versioned LLM verifier held constant across systems, and publishes every artifact. No vendor pays for inclusion, ranking, or removal.
What does the Structured Speech-to-Text Benchmark not measure?
Multilingual accuracy, diarization, timestamps, punctuation quality, noise/telephony robustness, domain vocabularies, custom vocabulary boosts, on-device deployment, and long-form behavior beyond the released clips. For your domain, run your own production clips. For raw Word Error Rate and STT latency, go to the Speech-to-Text benchmark.
