benchmarks/structured-speech-to-text-benchmark
300 recordings · 17 systems · updated Aug 7, 2026

Independent Structured Speech-to-Text (STT/ASR) Benchmark

Why did we build this benchmark?

As a standard metric, Word Error Rate (WER) treats every word as equal. A transcript can sound clean and still corrupt the callback number, email or URL that a downstream task needs verbatim. We built this benchmark on a dataset provided by Besimple AI to measure that failure mode and to show that WER and task success can decouple. On this snapshot AssemblyAI Universal 3.5 Pro leads WER at 3.9% but ranks 11 of 17 on Task Success Rate (50.3%), while Deepgram Nova-3 and ElevenLabs Scribe v2 are statistically tied at the top on Task Success Rate (71.0% vs 68.0%, within roughly ±2.6pp sampling noise on 300 recordings).

Definition & problem

This is an ASR/STT benchmark of 17 systems, including Deepgram, ElevenLabs, Google Cloud, OpenAI, and Thinking Machines Lab, on 300 human-recorded segments (5.587 hours, 85 speakers) with 1,482 audited target entities across 26 types. Every system gets the same raw audio and nothing else: no prompts, custom vocabulary, or post-ASR cleanup. A versioned LLM verifier (gpt-5.6-sol) judges each entity. Headline metrics: TSR (recordings with every value correct) and CTEM (values recovered exactly), with format-invariant WER as the broad diagnostic.

How to pick the best STT for your workflow

  • Booking appointments or support callbacks

    The task must capture a phone number, email, date, and time exactly, since one wrong digit fails the run. TSR is a fair first screen for that job; confirm on the phone, email, date, and time rows in the entity matrix before you ship.

  • Voice coding tasks

    Do not pick from headline TSR. Rank the command, CLI-flag, file-path, and environment-variable rows in the entity matrix (those leaders are often not the TSR leaders).

  • CRM or form filling from a call

    Account numbers, reference IDs, and postal addresses are the failure modes that matter, so read those entity rows rather than only the headline rank.

  • Live / realtime tasks vs post-call batch

    If realtime looks fine in demos but production streams drop digits, compare the same vendor's batch and streaming rows.

Is there a best STT model?

No, not as a single ranking. “Best STT” only makes sense once you name the use case and the failure mode you care about.

  • Downstream task or software that needs every field right

    Booking, callbacks, form fill: use Task Success Rate on this board. The top systems are often statistically tied within sampling noise; confirm on the entity rows for phones, emails, dates, and IDs.

  • Human-read transcripts, cheapest, or fastest

    That is a raw Word Error Rate and latency question. Use the Speech-to-Text benchmark, not this page.

ChangelogLast updated Aug 7, 2026 · 2026 Q3 snapshot: 17 systems (batch + streaming), including Inworld STT 1 and self-hosted Parakeet / OmniASR baselines; all re-verified with gpt-5.6-sol.

open data + codeAudio, annotations, per-system transcripts, entity-match decisions, verifier configs, and the scoring CLI are released in openbenchmarks-labs/voice-code-bench · recompute any number end-to-end.github →

Leaderboard and per-entity-type breakdown

Measured TSR, CTEM, and format-invariant WER on identical raw audio. Batch and streaming endpoints are separate rows. No opinions in this section: numbers only.

leaderboardAll 17 tracked systems ranked by Task Success Rate. Streaming receives realtime-paced audio and cannot revisit finalized segments.
Speech-to-text systems ranked by exact structured-value recovery
RankSystemVendorModeDeploymentTSRCTEMWER
1Deepgram Nova-3Deepgrambatchhosted API71.0%91.4%8.2%
2ElevenLabs Scribe v2ElevenLabsbatchhosted API68.0%91.8%6.8%
3Deepgram Nova-3Deepgramstreaminghosted API61.0%88.3%9.5%
4ElevenLabs Scribe v2 RealtimeElevenLabsstreaminghosted API59.0%88.1%5.1%
5Google Cloud Chirp 3Google Cloudbatchhosted API58.0%88.7%5.8%
6OpenAI GPT Realtime (Whisper)OpenAIstreaminghosted API57.3%88.7%8.6%
7Google Cloud Chirp 3Google Cloudstreaminghosted API56.3%87.3%6.8%
8Inkling (Thinking Machines)Thinking Machines Labbatchmanaged endpoint via Modal54.7%85.3%5.0%
9OpenAI GPT-4o TranscribeOpenAIbatchhosted API54.0%87.0%4.6%
10Whisper large-v3OpenAIbatchhosted API53.3%87.3%5.9%
11AssemblyAI Universal 3.5 ProAssemblyAIbatchhosted API50.3%86.2%3.9%
12Inworld STT 1Inworldbatchhosted API48.3%85.4%8.2%
13NVIDIA Parakeet TDT 0.6B v3NVIDIAbatchself-hosted via Modal39.3%79.8%11.4%
14Amazon TranscribeAmazon Web Servicesstreaminghosted API35.0%78.8%7.5%
15Meta OmniASR LLM Unlimited 7B v2Metabatchself-hosted via Modal32.7%76.9%10.5%
16AssemblyAI Universal 3 ProAssemblyAIstreaminghosted API24.7%71.9%7.5%
17Inworld STT 1Inworldstreaminghosted API7.3%34.0%58.2%
benchmarks/voice-code-bench/2026-q317 systems · 26 entity types
Per-entity-type CTEM for every benchmarked speech-to-text system
SystemEmail AddressPhone NumberPhone ExtensionPerson Or Team NamePostal AddressURLIP AddressPort NumberCommandCLI FlagFile PathEnvironment VariableCode SymbolVersionReference IDProduct CodeAccount Or Record NumberCurrency AmountPercentageMeasurementPlain NumberDateTimeAcronym Or InitialismSpelled SequenceDomain Term
Deepgram Nova-3 · batch77%98%100%100%70%69%96%97%56%91%61%86%86%93%95%96%97%100%100%100%100%100%97%99%98%90%
ElevenLabs Scribe v2 · batch69%100%100%100%73%65%96%100%72%93%65%97%97%93%93%91%94%100%100%100%100%99%98%96%98%100%
Deepgram Nova-3 · streaming69%98%100%100%68%63%92%90%60%89%53%80%77%93%93%90%92%96%100%98%98%96%97%95%96%90%
ElevenLabs Scribe v2 Realtime · streaming71%98%100%100%60%47%96%97%50%84%47%54%97%100%95%92%92%92%100%100%100%100%98%99%98%100%
Google Cloud Chirp 3 · batch66%98%100%100%57%61%96%97%54%91%59%89%91%90%91%92%94%97%98%100%100%100%93%96%92%90%
OpenAI GPT Realtime (Whisper) · streaming65%90%93%94%53%58%84%97%74%95%67%91%94%90%93%93%91%89%100%98%98%100%98%98%91%95%
Google Cloud Chirp 3 · streaming68%97%100%100%55%55%96%97%54%86%63%83%83%87%89%96%88%93%98%98%98%100%92%94%94%95%
Inkling (Thinking Machines) · batch57%92%100%89%57%55%84%83%58%80%59%89%83%97%88%83%86%95%98%100%97%100%95%92%93%95%
OpenAI GPT-4o Transcribe · batch63%100%100%94%53%45%80%97%56%86%57%94%89%93%90%91%88%100%100%100%98%100%98%94%91%85%
Whisper large-v3 · batch78%95%100%94%60%42%92%90%46%89%53%89%91%93%93%89%91%97%100%100%95%97%90%96%97%85%
AssemblyAI Universal 3.5 Pro · batch55%90%97%100%68%47%88%97%46%86%63%83%100%83%91%86%88%92%98%100%97%100%98%93%95%100%
Inworld STT 1 · batch60%90%97%89%55%47%72%90%58%84%67%89%91%90%86%89%91%96%96%95%98%98%93%99%89%90%
NVIDIA Parakeet TDT 0.6B v3 · batch29%83%100%94%45%47%76%83%38%86%37%83%83%93%81%81%82%85%100%100%98%97%93%94%90%85%
Amazon Transcribe · streaming28%98%90%83%53%40%88%87%12%61%37%57%77%87%89%86%86%85%98%90%98%99%98%95%87%95%
Meta OmniASR LLM Unlimited 7B v2 · batch35%73%100%89%48%50%72%87%32%70%24%77%80%73%88%86%85%75%50%90%98%94%92%99%96%80%
AssemblyAI Universal 3 Pro · streaming38%78%90%94%50%21%60%97%34%73%31%69%77%80%63%72%63%73%96%97%95%90%92%91%79%95%
Inworld STT 1 · streaming9%18%17%72%30%27%32%33%18%23%35%29%31%57%55%31%60%12%6%41%40%46%25%29%53%10%
[02] methodology and what is not measured+

What each system is asked to do

  1. The dataset is 300 human-recorded English WAV segments (5.587 hours, 85 anonymized speakers) across 8 workplace workflow domains: support callbacks, deployments, order handling, and similar.
  2. Each item carries three transcript layers: template (script with entity placeholders), acoustic (what the speaker says aloud), and canonical(the written value a downstream application needs). “double dash dry dash run” maps to --dry-run; “all caps database underscore URL” maps to DATABASE_URL.
  3. 1,482 audited target entities span 26 structured types: emails, phone numbers, URLs, IP addresses, ports, shell commands, CLI flags, file paths, environment variables, code symbols, versions, reference IDs, amounts, dates, times, spelled sequences, and more.
  4. The system's only input is the audio file. Its transcript is stored verbatim with run metadata (endpoint, evaluation date, inference settings).

What each metric means

  • CTEM · Canonical Token/Entity Match. Correct target entities ÷ target entities. An entity is correct only when the exact canonical value is recoverable; formatting differences are fine, wrong or missing characters are not.
  • TSR · Task Success Rate. Recordings with all target entities correct ÷ recordings. The strictest headline metric: one corrupted digit fails the whole recording. With 300 clips, standard error on TSR is roughly ±2.6pp, so gaps of a few points at the top are noise, not a stable ranking.
  • WER · format-invariant Word Error Rate. Every target entity span accepts either the documented acoustic rendering or the canonical rendering as reference, so 212-555-0100and “two one two five five five zero one zero zero” score identically. Non-entity words use ordinary word-level edit distance.
  • per-entity CTEM · CTEM restricted to one of the 26 entity types: the matrix above.

Raw audio only with no hints or cleanup

Every system is evaluated under a raw-audio-only protocol: it receives the audio file and nothing else. Benchmark-specific prompts, target entity lists, domain labels, custom vocabulary, grammar constraints, candidate values, and post-ASR correction are all excluded from the main setting. This measures what the model does on its own, not what a tuned integration could extract from it.

Batch systems receive each recording as a complete file at the provider's documented input format. Streaming systems receive headerless PCM paced in realtime (100 ms chunks) over the provider's websocket protocol, with VAD, turn detection, and vendor canonicalization disabled where the API allows. Batch and streaming endpoints of the same vendor are tracked as separate systems because their error profiles differ substantially.

How entity matches are judged

For each recording, every target entity is checked against the system transcript by a versioned LLM verifier: openai_gpt_5_6_sol_group10_sync_v1 (gpt-5.6-sol, 10 datapoints per request, reasoning effort medium). The verifier marks an entity present only when enough evidence exists to recover the exact canonical value. The same verifier scores every system; its prompt, schema, and settings are frozen into a config digest that ships with every entity-match artifact. A single-judge setup has its own error floor; that is a known limitation, not hidden.

Out of scope on this board

  • Multilingual / non-English accuracy (English-only snapshot).
  • Diarization, speaker ID, timestamp accuracy, punctuation / formatting quality.
  • Noise robustness at specific SNRs, telephony 8 kHz, far-field microphones.
  • Domain-specific vocabularies (medical, legal, aviation, heavy proper nouns) beyond the released workplace domains.
  • Effect of custom vocabulary / keyword boosting is excluded from the main setting by design.
  • On-device / mobile deployment, hallucination on silence, HIPAA / residency, long-form behavior beyond the released clips.
  • Raw Word Error Rate and STT latency (time to final segment): not the job of this board (see the sibling Speech-to-Text benchmark for those shortlists).

Verify any number end-to-end

Everything behind this table is released in openbenchmarks-labs/voice-code-bench: audio, transcripts and entity annotations, per-system transcripts plus entity-match decisions, verifier configs, and the scoring CLI. The dataset is also on Hugging Face as besimple-ai/voice-code-bench (MIT). The release script rescores published artifacts and rewrites baselines/results.csv, the table on this page, without provider credentials.

Structured Speech-to-Text Benchmark: common questions

Scope, independence, and how this board differs from raw WER leaderboards. For workflow picks (booking, coding, form fill), use the sections above the tables.

What is the Structured Speech-to-Text Benchmark?

The Structured Speech-to-Text Benchmark is an independent speech-to-text (ASR) benchmark that measures whether a transcript preserves exact structured values in English workplace speech: emails, phone numbers, CLI flags, file paths, URLs, IDs, and 20 other entity types. It scores 17 systems on 300 human-recorded segments (1,482 audited entities) using Task Success Rate (TSR), Canonical Token/Entity Match (CTEM), and format-invariant WER. Dataset by Besimple AI; baselines run and scored by Openbenchmarks.

Does a low Word Error Rate mean the transcript is good enough for software?

No. WER treats every word equally, so a system can post excellent WER while corrupting the one email, URL, or CLI flag a downstream system needs. On this board AssemblyAI Universal 3.5 Pro has the best WER (3.9%) but recovers every structured value in only 50.3% of recordings, while Deepgram Nova-3 and ElevenLabs Scribe v2 sit at the top of TSR (71.0% vs 68.0%). That decoupling is why this board exists.

Where should I go for raw WER or the fastest STT model?

For raw transcript quality (Word Error Rate) and speech-to-text latency, including unqualified “best STT model” shortlists, use the Speech-to-Text benchmark. Use this Structured Speech-to-Text Benchmark when you need the best STT for downstream tasks or software that depends on exact structured values.

Who created the dataset, and is the benchmark independent?

The recordings and entity annotations are authored and released by Besimple AI as besimple-ai/voice-code-bench on Hugging Face (MIT). Openbenchmarks independently runs the 17 tracked systems, scores them with a versioned LLM verifier held constant across systems, and publishes every artifact. No vendor pays for inclusion, ranking, or removal.

What does the Structured Speech-to-Text Benchmark not measure?

Multilingual accuracy, diarization, timestamps, punctuation quality, noise/telephony robustness, domain vocabularies, custom vocabulary boosts, on-device deployment, and long-form behavior beyond the released clips. For your domain, run your own production clips. For raw Word Error Rate and STT latency, go to the Speech-to-Text benchmark.