Openbenchmarks
Independent benchmarks for build-vs-buy decisions. Verifiable, reproducible and open source.
Last run on 28 Sept 2026Benchmarks
Web Search for AI
Web Search APIs benchmarked across three tasks: factual lookup, hard retrieval, and multi-hop search.
- Most accurate
Exa fast99.3%
- Fastest
Parallel turbo348 ms- Lowest cost
TinyFishFree
- Most tasks
Perplexity77.3%- Fastest
Perplexity18.3 s- Lowest cost
TinyFish$0.035
- Most tasks
Exa deep83.0%
- Fastest
Perplexity21.5 s- Lowest cost
TinyFish$0.050
- Leads F1
Parallel basic46.5- Fastest
Brave Search43.5 s- Lowest cost
TinyFish$0.212
- Leads F1
Exa deep48.2
- Fastest
Brave Search45.1 s- Lowest cost
TinyFish$0.219
Multi Turn Company Search
Search APIs tested as tools for a research agent finding complete company sets across three- and four-part constraints. Ranked on precision, recall, F1, exact-set accuracy, latency, and cost.
- Leads F1
Parallel basic46.5- Fastest
Brave Search43.5 s- Lowest cost
TinyFish$0.212
- Leads F1
Exa deep48.2
- Fastest
Brave Search45.1 s- Lowest cost
TinyFish$0.219
Company News
Which API answers questions about recent company events. Web search APIs against dedicated news indexes, same questions.
- Most accurate
Exa (type=fast)99.3%
- Fastest
Parallel (mode=turbo)348 ms- Lowest cost
TinyFishFree
Web Search for Coding Agents
Which search API helps a coding agent find an opaque detail in official Enterprise SaaS docs and ground the patch in a URL it actually retrieved.
- Most tasks
Perplexity77.3%- Fastest
Perplexity18.3 s- Lowest cost
TinyFish$0.035
- Most tasks
Exa deep83.0%
- Fastest
Perplexity21.5 s- Lowest cost
TinyFish$0.050
Document Processing Benchmark
Specialised document parsers benchmarked on heavily redlined contracts, delivered as text-layer PDFs and as scanned images.
- Most accurate
LlamaParse agentic80.0%- Fastest
Reducto7s- Parser $ / 1k correct
Mistral OCR$14.04
- Most accurate
Datalab track changes79.7%- Fastest
Reducto8s- Parser $ / 1k correct
Datalab track changes$19.13
Company Lookalikes
Tests how well company-search and lookalike APIs turn a seed company domain or description into a ranked list of genuinely similar businesses. Use it to compare tools for account discovery, prospecting, market mapping, and TAM expansion.
Company Firmographic Enrichment
282 company domains × 11 APIs, normalized into the same seven-field active scoring contract. Compare enrichment success rate, field accuracy, coverage, company match rate, latency, and cost by workflow.
Company Funding Data Enrichment
17 funding vendors, judged on latest funding stage against human labelled ground truth.
Web Search for AI
Web Search APIs benchmarked across three tasks: factual lookup, hard retrieval, and multi-hop search.
Company News
Which API answers questions about recent company events. Web search APIs against dedicated news indexes, same questions. Ranked by $ / 1k correct.
Multi Turn Company Search
Search APIs tested as tools for a research agent finding complete company sets across three- and four-part constraints. Ranked on precision, recall, F1, exact-set accuracy, latency, and cost.
Web Search for Coding Agents
Which search API helps a coding agent find an opaque detail in official Enterprise SaaS docs and ground the patch in a URL it actually retrieved.
Inference Benchmark
Serverless inference providers serve the same GLM 5.3 Flash model on 600 latency-sensitive lookup, classification, and extraction requests each. Compare response latency, streaming speed, task success, and failures without collapsing them into one overall rank.
Voice Agents Latency Benchmark
TTFAB — how long a caller waits before a voice AI agent starts speaking — measured over real phone calls from the call's own audio, never platform-reported timestamps.
Text-to-Speech
31 text-to-speech models on Time to First Audio (TTFA) + Word Error Rate, measured by Coval under production-realistic conditions. Mirrored with attribution.
Live Speech-to-Text
29 live speech-to-text models on Word Error Rate and Time to Final Segment (TTFS), measured by Coval under production-realistic conditions. Rolling 7-day window, mirrored with attribution. Ranked here by WER, not median TTFS.
Structured Speech-to-Text
300 human-recorded workplace utterances × 17 ASR systems, checked for exact recovery of 1,482 structured values — emails, phone numbers, CLI flags, file paths, IDs. Cell value is Task Success Rate — recordings with every value correct.