benchmarks/web-search/most accurate for LLM apps
open source code · open data · 3 tasks · accuracy only

Most accurate web search API for LLM apps

Most accurate, by audience: AI and LLM agents → · Most accurate web search API for LLM agents

Independent 2026 benchmark (open source code, open data). Accuracy of web search APIs for AI agents and LLM agents on each of the three tasks. This page keeps accuracy columns only.

How to pick. Start from the task that matches your workflow, then read that accuracy board.

  • One fact, one query. Read factual lookup. Prefer Accuracy, then AR@1 or AR@5.
  • Seated deep in a unique page, across very similar pages. Read hard retrieval. Prefer Task completion on the search & fetch or search-only board that matches your tool setup.
  • One query is not enough. Read multi-hop search. Prefer F1, then precision versus recall.

Full web search benchmark →

Factual lookup

Extracted-answer accuracy, AR@1, and AR@5 on identical queries, max 10 results.

View full factual lookup benchmark: company news

Factual lookup: web search APIs ranked by extracted-answer accuracy
RankVendorEndpoint & configurationAccuracyAR@1AR@5
1Exaweb searchPOST /searchtype=fast99.3%95.0%99.3%
2Exaweb searchPOST /searchtype=instant97.7%80.0%97.3%
3Perplexityweb searchPOST /searchsearch_context_size=low97.3%91.7%98.3%
4Linkupweb searchPOST /v1/searchdepth=fast · outputType=searchResults96.7%79.0%94.7%
5SERPgoogle searchGET google-search74.p.rapidapi.comlimit=1096.0%78.0%95.0%
6Firecrawlweb searchPOST /v2/search95.3%77.7%96.7%
7Brave Searchweb searchPOST /res/v1/llm/contextcount=1094.0%81.0%94.7%
8Youweb searchPOST /v1/searchcount=1093.3%84.7%94.7%
9Parallelweb searchPOST /v1/searchmode=basic93.3%55.0%92.0%
10Brave Searchweb searchGET /res/v1/web/searchcount=10 · result_filter=web93.3%79.3%91.7%
11TinyFishweb searchGET api.search.tinyfish.aifree · 30 req/min cap92.0%74.3%90.7%
12Linkupweb searchPOST /v1/searchdepth=standard · outputType=searchResults92.0%67.3%90.3%
13Youweb searchPOST /v1/searchextraction_mode=highlights90.7%72.0%89.7%
14Parallelweb searchPOST /v1/searchmode=fast86.0%44.3%79.0%
15Parallelweb searchPOST /v1/searchmode=turbo71.3%45.3%66.0%
16Tavilyweb searchPOST /searchsearch_depth=ultra-fast13.3%10.0%17.0%

Hard retrieval

Grounded task completion. Search & fetch and search-only are separate rankings.

View full hard retrieval benchmark: web search for coding agents

[02.a]search & fetch
Hard retrieval, search & fetch: ranked by task completion
RankVendorEndpoint & configurationTask completion
1Exa deepPOST /search type=deepPOST /contents83.0 ± 1.0
2Exa autoPOST /search type=autoPOST /contents81.7 ± 1.1
3TinyFishGET api.search.tinyfish.aiGET api.fetch.tinyfish.ai format=markdown79.0 ± 2.0
4PerplexityPOST /search search_context_size=highPOST /search search_context_size=high77.7 ± 1.5
5Parallel advancedPOST /v1/search mode=advancedPOST /v1/extract77.0 ± 1.0
6Parallel basicPOST /v1/search mode=basicPOST /v1/extract76.0 ± 0.0
7FirecrawlPOST /v2/searchPOST /v2/scrape76.0 ± 1.0
8YouPOST /v1/searchPOST /v1/contents formats=markdown61.7 ± 0.6
9Tavily advancedPOST /search search_depth=advancedPOST /extract extract_depth=advanced60.0 ± 2.0
10Tavily basicPOST /search search_depth=basicPOST /extract extract_depth=basic59.0 ± 1.7
11Linkup standardPOST /v1/search depth=standardPOST /v1/fetch mode=standard48.3 ± 4.9

Task completion is mean ± SD of 3 runs; n = 100 tasks.

[02.b]search-only
Hard retrieval, search-only: ranked by task completion
RankVendorEndpoint & configurationTask completion
1PerplexityPOST /search search_context_size=low77.3 ± 2.1
2FirecrawlPOST /v2/search70.3 ± 1.5
3Parallel fastPOST /v1/search mode=fast66.7 ± 1.5
4Exa fastPOST /search type=fast66.3 ± 1.5
5Parallel turboPOST /v1/search mode=turbo64.7 ± 2.1
6Exa instantPOST /search type=instant61.3 ± 2.9
7TinyFishGET api.search.tinyfish.ai59.3 ± 1.5
8Tavily fastPOST /search search_depth=fast47.3 ± 2.5
9Linkup fastPOST /v1/search depth=fast43.3 ± 1.1
10BravePOST /res/v1/llm/context43.0 ± 2.0
11YouPOST /v1/search39.3 ± 2.5

Task completion is mean ± SD of 3 runs; n = 100 tasks.

Multi-hop search

F1, precision, recall, and exact-set accuracy. Search-only and search & fetch are separate rankings.

View full multi-hop search benchmark: multi-turn company search

Web Search only

Multi-hop, search-only: ranked by F1
ProviderEndpoint & configurationF1PrecisionRecallExact set
ParallelPOST /v1/search mode=basicparallel-basic46.5 ± 1.988.7 ± 0.934.4 ± 2.03.7 ± 2.6
ExaPOST /search type=deepexa-deep45.4 ± 2.083.2 ± 3.033.7 ± 1.41.5 ± 1.3
ParallelPOST /v1/search mode=advancedparallel-advanced44.2 ± 1.487.6 ± 3.432.0 ± 1.62.2 ± 0.0
ExaPOST /search type=instantexa-instant43.3 ± 1.082.6 ± 3.832.2 ± 0.53.0 ± 1.3
LinkupPOST /v1/search depth=fastlinkup-fast41.1 ± 1.782.8 ± 1.030.3 ± 1.50.7 ± 1.3
TavilyPOST /search search_depth=advancedtavily-advanced41.1 ± 2.383.7 ± 4.229.8 ± 1.72.2 ± 0.0
LinkupPOST /v1/search depth=standardlinkup-standard40.6 ± 0.984.0 ± 7.929.7 ± 1.41.5 ± 2.6
ParallelPOST /v1/search mode=fastparallel-fast38.0 ± 2.079.9 ± 4.627.5 ± 1.41.5 ± 1.3
PerplexityPOST /searchsearch_context_size=low37.8 ± 2.179.3 ± 5.726.8 ± 1.42.2 ± 2.2
ParallelPOST /v1/search mode=turboparallel-turbo34.7 ± 2.480.0 ± 5.224.8 ± 2.41.5 ± 1.3
YouPOST /v1/searchyou33.1 ± 2.475.7 ± 3.123.1 ± 1.91.5 ± 1.3
FirecrawlPOST /v2/searchfirecrawl30.4 ± 1.177.3 ± 4.420.7 ± 0.72.2 ± 0.0
Brave SearchGET /res/v1/web/searchbrave28.0 ± 1.766.9 ± 8.519.3 ± 0.50.7 ± 1.3
TinyFishGET api.search.tinyfish.aitinyfish26.6 ± 1.364.7 ± 1.517.9 ± 0.81.5 ± 1.3
SeltzPOST /v1/search scope=companiesseltz-companies14.5 ± 0.940.0 ± 3.19.4 ± 0.70.0 ± 0.0
SERP (RapidAPI)GET google-search74.p.rapidapi.comserp0.4 ± 0.60.7 ± 1.30.3 ± 0.40.0 ± 0.0

F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points.

Web Search + Fetch

Multi-hop, search & fetch: ranked by F1
ProviderEndpoint & configurationF1PrecisionRecallExact set
ExaPOST /search type=deepexa-deep48.2 ± 2.189.4 ± 1.236.0 ± 2.22.2 ± 2.2
PerplexityPOST /searchsearch_context_size=high46.6 ± 2.087.7 ± 5.934.7 ± 1.12.2 ± 2.2
ExaPOST /search type=instantexa-instant44.9 ± 0.985.9 ± 2.533.5 ± 0.95.2 ± 1.3
ParallelPOST /v1/search mode=basicparallel-basic42.3 ± 1.181.3 ± 2.831.3 ± 1.13.0 ± 1.3
ParallelPOST /v1/search mode=advancedparallel-advanced42.2 ± 1.187.6 ± 0.330.1 ± 1.12.2 ± 0.0
LinkupPOST /v1/search depth=standardlinkup-standard42.0 ± 1.890.7 ± 2.030.5 ± 2.03.0 ± 3.4
TavilyPOST /search search_depth=advancedtavily-advanced41.0 ± 1.389.4 ± 6.029.1 ± 0.72.2 ± 0.0
LinkupPOST /v1/search depth=fastlinkup-fast39.9 ± 1.385.3 ± 3.628.6 ± 1.30.7 ± 1.3
ParallelPOST /v1/search mode=fastparallel-fast39.3 ± 3.382.3 ± 6.628.2 ± 2.02.2 ± 0.0
ParallelPOST /v1/search mode=turboparallel-turbo36.0 ± 3.583.6 ± 4.025.0 ± 2.60.0 ± 0.0
YouPOST /v1/searchyou34.0 ± 0.978.8 ± 0.623.8 ± 1.23.0 ± 1.3
FirecrawlPOST /v2/searchfirecrawl33.2 ± 2.183.3 ± 0.822.7 ± 1.81.5 ± 1.3
TinyFishGET api.search.tinyfish.aitinyfish30.2 ± 3.570.9 ± 11.320.9 ± 2.50.0 ± 0.0
Brave SearchGET /res/v1/web/searchbrave29.4 ± 1.673.5 ± 5.820.4 ± 0.91.5 ± 1.3
SeltzPOST /v1/search scope=companiesseltz-companies16.3 ± 1.549.5 ± 2.410.2 ± 1.20.0 ± 0.0
SERP (RapidAPI)GET google-search74.p.rapidapi.comserp0.0 ± 0.00.0 ± 0.00.0 ± 0.00.0 ± 0.0

F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points.

Most accurate web search API for AI and LLM agents: FAQ

What is the most accurate web search API for AI agents?

There is no single most accurate web search API for AI agents across tasks. This 2026 benchmark ranks the same vendors on three jobs — factual lookup, hard retrieval, and multi-hop search — and keeps the rankings separate. On factual lookup, Exa (type=fast) leads extracted-answer accuracy at 99.3%. On hard retrieval (search & fetch), Exa deep leads task completion at 83.0%. On multi-hop (search-only), Parallel basic leads F1 at 46.5%.

What is the most accurate web search API for LLM agents?

Same board: the most accurate web search API for LLM agents depends on the task. Rankings for factual lookup, hard retrieval, and multi-hop search are not averaged. The model is held constant so the comparison isolates the search API.

How is accuracy measured on each task?

Factual lookup (300 questions): extracted-answer accuracy, AR@1, AR@5. Hard retrieval (100 documentation tickets): grounded task completion, search & fetch and search-only separately. Multi-hop (45 company-discovery questions): F1, precision, recall, and exact-set accuracy, search-only and search & fetch separately.

Where are the fastest, cheapest, and most token-efficient rankings?

This page is accuracy only. Fastest · Cheapest · Most token-efficient.