Backed by Y Combinator

Openbenchmarks

Independent benchmarks for build-vs-buy decisions. Verifiable, reproducible and open source.

Last run on 28 Sept 2026

Benchmarks

Web Search for AI

Web Search APIs benchmarked across three tasks: factual lookup, hard retrieval, and multi-hop search.

factual lookup
Most accurate
Exa fast99.3%
Fastest
Parallel turbo348 ms
Lowest cost
TinyFishFree
hard retrieval · search only
Most tasks
Perplexity77.3%
Fastest
Perplexity18.3 s
Lowest cost
TinyFish$0.035
hard retrieval · search + fetch
Most tasks
Exa deep83.0%
Fastest
Perplexity21.5 s
Lowest cost
TinyFish$0.050
multi-hop · search only
Leads F1
Parallel basic46.5
Fastest
Brave Search43.5 s
Lowest cost
TinyFish$0.212
multi-hop · search + fetch
Leads F1
Exa deep48.2
Fastest
Brave Search45.1 s
Lowest cost
TinyFish$0.219
open benchmark→

Multi Turn Company Search

Search APIs tested as tools for a research agent finding complete company sets across three- and four-part constraints. Ranked on precision, recall, F1, exact-set accuracy, latency, and cost.

search only
Leads F1
Parallel basic46.5
Fastest
Brave Search43.5 s
Lowest cost
TinyFish$0.212
search + fetch
Leads F1
Exa deep48.2
Fastest
Brave Search45.1 s
Lowest cost
TinyFish$0.219
open benchmark→

Company News

Which API answers questions about recent company events. Web search APIs against dedicated news indexes, same questions.

Most accurate
Exa (type=fast)99.3%
Fastest
Parallel (mode=turbo)348 ms
Lowest cost
TinyFishFree
open benchmark→

Web Search for Coding Agents

Which search API helps a coding agent find an opaque detail in official Enterprise SaaS docs and ground the patch in a URL it actually retrieved.

search only
Most tasks
Perplexity77.3%
Fastest
Perplexity18.3 s
Lowest cost
TinyFish$0.035
search + fetch
Most tasks
Exa deep83.0%
Fastest
Perplexity21.5 s
Lowest cost
TinyFish$0.050
open benchmark→

Document Processing Benchmark

Specialised document parsers benchmarked on heavily redlined contracts, delivered as text-layer PDFs and as scanned images.

text-layer PDF
Most accurate
LlamaParse agentic80.0%
Fastest
Reducto7s
Parser $ / 1k correct
Mistral OCR$14.04
scanned PDF · image only
Most accurate
Datalab track changes79.7%
Fastest
Reducto8s
Parser $ / 1k correct
Datalab track changes$19.13
open benchmark→

Company Lookalikes

Tests how well company-search and lookalike APIs turn a seed company domain or description into a ranked list of genuinely similar businesses. Use it to compare tools for account discovery, prospecting, market mapping, and TAM expansion.

ocean.ioexaparallelpredictleads
view methodology→
open data + code · github.com/openbenchmarks-labs/lookalikes

Company Firmographic Enrichment

282 company domains × 11 APIs, normalized into the same seven-field active scoring contract. Compare enrichment success rate, field accuracy, coverage, company match rate, latency, and cost by workflow.

Correct field yield · higher is better · scale 0–100%
People Data Labs89.0%
Apollo88.6%
Parallel86.0%
Predict Leads80.6%
Explorium70.5%
+6 more
open benchmark→
open data + code · github.com/openbenchmarks-labs/company-enrichment

Company Funding Data Enrichment

17 funding vendors, judged on latest funding stage against human labelled ground truth.

Correct stage yield · higher is better · scale 0–100%
Firecrawl92.3%
Parallel90.0%
Parallel Responses API (medium)89.0%
Exa88.7%
Exa88.6%
+18 more
open benchmark→
open data + code · github.com/openbenchmarks-labs/company-funding

Web Search for AI

Web Search APIs benchmarked across three tasks: factual lookup, hard retrieval, and multi-hop search.

Multi-hop search F1 · higher is better · scale 0–100%
Parallel basic47%
Exa deep45%
Parallel advanced44%
Exa instant43%
Linkup fast41%
+15 more
open benchmark→
open code · github.com/openbenchmarks-labs/multi-turn-company-search

Company News

Which API answers questions about recent company events. Web search APIs against dedicated news indexes, same questions. Ranked by $ / 1k correct.

$ / 1k correct · lower is better
Parallel fast$1.16
Parallel turbo$1.40
Nimble lite$1.46
SERP$3.13
Perplexity$5.14
+17 more
open benchmark→

Multi Turn Company Search

Search APIs tested as tools for a research agent finding complete company sets across three- and four-part constraints. Ranked on precision, recall, F1, exact-set accuracy, latency, and cost.

Search-only F1 · higher is better · scale 0–100%
Parallel basic47%
Exa deep45%
Parallel advanced44%
Exa instant43%
Linkup fast41%
+15 more
open benchmark→

Web Search for Coding Agents

Which search API helps a coding agent find an opaque detail in official Enterprise SaaS docs and ground the patch in a URL it actually retrieved.

Task completion · higher is better · scale 0–100%
Perplexity77%
Firecrawl70%
Parallel fast67%
Exa fast66%
Parallel turbo65%
+9 more
open benchmark→

Inference Benchmark

Serverless inference providers serve the same GLM 5.3 Flash model on 600 latency-sensitive lookup, classification, and extraction requests each. Compare response latency, streaming speed, task success, and failures without collapsing them into one overall rank.

E2E latency P99 · lower is better
Baseten3229 ms
Modal3279 ms
Telnyx3478 ms
Novita AI7248 ms
Z.AI8591 ms
+5 more
open benchmark→
open runner · github.com/openbenchmarks-labs/inference

Voice Agents Latency Benchmark

TTFAB — how long a caller waits before a voice AI agent starts speaking — measured over real phone calls from the call's own audio, never platform-reported timestamps.

Telnyx1296 ms
ElevenLabs1424 ms
Bland AI1520 ms
Vapi1558 ms
Retell AI1740 ms
p50 p90 · p95 p992,078 turns measured
open benchmark→
open data + code · github.com/openbenchmarks-labs/voice-agent-latency

Text-to-Speech

31 text-to-speech models on Time to First Audio (TTFA) + Word Error Rate, measured by Coval under production-realistic conditions. Mirrored with attribution.

Median TTFA · lower is better
vui50 ms
gradium-tts-beta-20260951 ms
gradium-tts-beta62 ms
inworld-tts-2-flash64 ms
qwen3-tts-fast67 ms
+26 more
open benchmark→

Live Speech-to-Text

29 live speech-to-text models on Word Error Rate and Time to Final Segment (TTFS), measured by Coval under production-realistic conditions. Rolling 7-day window, mirrored with attribution. Ranked here by WER, not median TTFS.

Word Error Rate · lower is better · scale 0–100%
universal-3.5-pro2.47%
qwen3-asr-fast2.72%
realtime2.76%
qwen3-asr-1.7b2.78%
gemini-3.5-transcribe-live3.01%
+24 more
open benchmark→

Structured Speech-to-Text

300 human-recorded workplace utterances × 17 ASR systems, checked for exact recovery of 1,482 structured values — emails, phone numbers, CLI flags, file paths, IDs. Cell value is Task Success Rate — recordings with every value correct.

Task Success Rate · higher is better · scale 0–100%
Deepgram Nova-371%
ElevenLabs Scribe v268%
Deepgram Nova-3 (streaming)61%
ElevenLabs Scribe v2 Realtime (streaming)59%
Google Cloud Chirp 358%
+12 more
open benchmark→