benchmarks/web-search
3 tasks · 12 vendors · last measured 26 Sept 2026

Web Search Benchmark

Exa (type=fast) leads factual lookup at 99.3% across 300 questions. Exa deep leads hard retrieval at 83.0% on 100 tickets. Parallel basic leads multi-hop search at 46.5% F1 on 45 questions. Independent benchmark of web search APIs for AI agents across 3 tasks. Open source. 300 lookups, 100 hard retrieval, 45 deep research. Last measured 26 Sept 2026.

Why we built this. The best web search API for AI agents and LLM agents depends on the job. A lookup that answers one fact is different from a retrieval use case that has to find answers buried in pages, which is different from a multi-hop search. There is no single best web search API — the winner changes with the tool mode too: Exa deep leads hard retrieval when the agent can fetch pages, Perplexity when it can only read snippets, and on multi-hop it is Parallel basic search-only against Exa deep with fetch. This 2026 board measures the same providers on those three tasks.

Model-only baseline: 0.The model cannot answer any question on this benchmark without web search. That is how the tasks were created: a correct answer has to be found, not recalled. Factual lookup uses latest company facts dated after the model's cutoff. Hard retrieval only passes if the patch cites a ground source URL from that run's search results; without a search tool the model cannot provide one. Multi-hop questions combine three or four constraints that are not conclusively present in training data; we confirmed that in a run with no search tool, where the model still scored 0.

How it is measured.We keep the model and judge constant for all search APIs in each task so the comparison isolates the web search API's performance. Because the model-only score is 0, the ranking is the search API, not the model's memory.

Factual lookup

A factual lookup is a single question with one verifiable answer. Search either returns that fact or it does not. Every question is a recent company event dated after the model's cutoff, so the model cannot answer it from memory.

Providers are listed alphabetically. Compare accuracy, answer recall, latency, token usage, and cost side by side, with no headline metric or overall performance ranking.

View full factual lookup benchmark: company news →

Exa fast is most accurate at 99.3%Parallel turbo is fastest at 348 msTinyFish costs least per correct answer at Free

Top 8 per metric
Factual lookup: web search APIs listed alphabetically on identical queries.
VendorEndpoint & configurationAccuracyAR@1AR@5$ / 1k correctLatencySnippet tokensTotal $Official docsCost per 1,000 queries
Brave Searchweb searchPOST /res/v1/llm/contextcount=1094.0%81.0%94.7%$5.32601ms2,064$1.50Official docs$5 / 1kSearch plan · LLM Context
Brave Searchweb searchGET /res/v1/web/searchcount=10 · result_filter=web93.3%79.3%91.7%$5.36630ms817$1.50Official docs$5 / 1k
Exaweb searchPOST /searchtype=fast99.3%95.0%99.3%$7.05652ms1,987$2.10Official docs$7 / 1ktype=fast · up to 10 results
Exaweb searchPOST /searchtype=instant97.7%80.0%97.3%$7.17398ms2,128$2.10Official docs$7 / 1ktype=instant · up to 10 results
Firecrawlweb searchPOST /v2/search95.3%77.7%96.7%$5.24510ms678$1.50Official docs$5 / 1k2 credits / 10 results
Linkupweb searchPOST /v1/searchdepth=fast · outputType=searchResults96.7%79.0%94.7%$5.171.57s3,022$1.50Official docs$5 / 1kdepth=fast · searchResults
Linkupweb searchPOST /v1/searchdepth=standard · outputType=searchResults92.0%67.3%90.3%$5.432.55s2,983$1.50Official docs$5 / 1kdepth=standard · searchResults
Nimbleweb searchPOST /v2/searchsearch_depth=lite · full_content=false · focus=general75.3%64.0%75.3%$1.463.15s413$0.33Official docs$1.1 / 1ksearch_depth=lite · full_content=false
Nimbleweb searchPOST /v2/searchsearch_depth=standard · full_content=false · focus=general93.0%86.3%94.3%$5.38861ms2,730$1.50Official docs$5 / 1ksearch_depth=standard · full_content=false
Parallelweb searchPOST /v1/searchmode=basic93.3%55.0%92.0%$5.361.68s2,330$1.50Official docs$5 / 1kmode=basic · 10 results
Parallelweb searchPOST /v1/searchmode=fast86.0%44.3%79.0%$1.16942ms1,839$0.30Official docs$1 / 1kmode=fast · 10 results
Parallelweb searchPOST /v1/searchmode=turbo71.3%45.3%66.0%$1.40348ms1,853$0.30Official docs$1 / 1kmode=turbo · 10 results
Perplexityweb searchPOST /searchsearch_context_size=low97.3%91.7%98.3%$5.141.38s476$1.50Official docs$5 / 1kSearch API · POST /search
SERPgoogle searchGET google-search74.p.rapidapi.comlimit=1096.0%78.0%95.0%$3.13751ms497$0.90Official docs$3 / 1kPro overage $0.003/request
Tavilyweb searchPOST /searchsearch_depth=advanced93.0%65.0%92.7%$17.204.29s2,210$4.80Official docs$16 / 1k2 credits · $0.008 PAYG
Tavilyweb searchPOST /searchsearch_depth=basic87.7%82.0%87.7%$9.131.88s1,639$2.40Official docs$8 / 1k1 credit · $0.008 PAYG
TinyFishweb searchGET api.search.tinyfish.aifree · 30 req/min cap92.0%74.3%90.7%Free30 req/min cap2.62s441FreeOfficial docs$0free · 30 req/min, 500/hour
Youweb searchPOST /v1/searchextraction_mode=highlights90.7%72.0%89.7%$5.51628ms2,837$1.50Official docs$5 / 1khighlights included in Web Search
Youweb searchPOST /v1/searchextraction_mode=highlights · knowledge=core92.0%67.3%89.7%$5.43889ms2,862$1.50Official docs$5 / 1khighlights · knowledge=core · same Web Search price assumption

$ / 1k correct is $ / 1k queries divided by extracted-answer accuracy. Providers are listed alphabetically; no single metric determines their order. Cost per 1,000 queries is the published PAYG list price, linked to the vendor pricing page. Not promotional packs or volume discounts. TinyFish Search is free with a 30 requests/min cap — $0 is not unlimited throughput.

Open CodeHarness, scoring and vendor runners in openbenchmarks-labs/factual-lookup-company-news-search.github →public dataset100 public company-news questions and ground truth answers on Hugging Face as openbenchmarks/OB-News-Websearch.huggingface →
[01.a] methodology and metric definitions+

How the lookup task is built

We measure the web search API. Every endpoint answers the same 300 lookup questions. The query, the model that reads the results and extracts an answer (gpt-5.6-terra, medium), and the judge (claude-opus-5 via Amazon Bedrock) are held constant. The only thing that changes is which search API is called.

  1. Collect official wires and newsrooms. Fix ground truth by human labelling before calling any vendors. Set of 300 company-news questions in v1.
  2. Send every endpoint the same natural-language question, one request, max 10 results. The query is the user question, unchanged. Nothing is rewritten per vendor.
  3. Persist the raw HTTP envelope, normalized hits (url, title, snippet), and latency for every call.
  4. An LLM in the harness (gpt-5.6-terra, medium) reads title + excerpt only and extracts an answer. A separate post-hoc judge from an independent model family (claude-opus-5 via Amazon Bedrock) scores accuracy and AR@K against the human labelled ground truth. It never fetches the live page.
  5. Score accuracy, AR@1, AR@5, snippet tokens, latency, and list-price cost. Providers are listed alphabetically.

A 100-question public set is on Hugging Face as openbenchmarks/OB-News-Websearch. Board scores use the locked 300-question set. Harness, scoring and vendor runners are in openbenchmarks-labs/factual-lookup-company-news-search.

What each metric means

  • Accuracy · gpt-5.6-terra (medium) writes an answer from the returned titles and snippets. The judge scores that written answer against human labelled ground truth.
  • AR@K · answer recall at K. Whether any of the top K vendor snippets already contained the human labelled ground truth. No extract.
  • AR@1 · the first snippet already contained the human labelled ground truth.
  • AR@5 · whether any one of the top 5 snippets already contained the full human labelled ground truth.
  • Latency · mean wall time of the search request, not the extract or score calls.
  • Total $ · search list price for the run. Judge tokens are not in this column.
  • Cost per 1,000 queries · published PAYG list price of the endpoint we called, linked to the vendor pricing page. Per-request search APIs are $ / 1,000 queries. Autobound and Datahyena bill credits, so that cell is $ / 1,000 credits.
  • $ / 1k correct · $ / 1k queries divided by extracted-answer accuracy. Same cost-to-quality figure as the cheapest page.

Hard retrieval

Hard retrieval is finding an answer that is seated deep within a unique page, across very similar pages. The snippet is not enough, and the model cannot guess the detail. Both boards: search & fetch, then search-only.

View full hard retrieval benchmark: web search for coding agents →

[02.a]search & fetch

Exa deep completes most tasks at 83.0%Perplexity is fastest at 21.5 sTinyFish costs least at $0.050

Top 8 per metric
Search & fetch coding-agent board, listed alphabetically
VendorEndpoint & configurationTask completionMedian task timeMedian task costAvg search timeMedian task tokensAPI list price
Exa autoPOST /search type=autoPOST /contents81.7 ± 1.123s$0.107$0.032 search+fetch$0.074 token cost1.19s27,433$0.007 / search$0.001 / fetch
Exa deepPOST /search type=deepPOST /contents83.0 ± 1.037s$0.127$0.057 search+fetch$0.070 token cost3.97s23,660$0.012 / search$0.001 / fetch
FirecrawlPOST /v2/searchPOST /v2/scrape76.0 ± 1.034s$0.082$0.028 search+fetch$0.055 token cost2.81s17,379$0.005 / search$0.0025 / fetch
Linkup standardPOST /v1/search depth=standardPOST /v1/fetch mode=standard48.3 ± 4.948s$0.147$0.038 search+fetch$0.108 token cost2.00s57,791$0.005 / search$0.005 / fetch
NimblePOST /v2/search search_depth=lite · full_content=false · focus=generalPOST /v2/extract formats=[markdown]60.3 ± 2.540s$0.054$0.054 token cost1.93s16,828$0.0011 / search
NimblePOST /v2/search search_depth=standard · full_content=false · focus=generalPOST /v2/extract formats=[markdown]45.0 ± 1.044s$0.078$0.078 token cost905ms28,586$0.005 / search
Parallel advancedPOST /v1/search mode=advancedPOST /v1/extract77.0 ± 1.033s$0.096$0.019 search+fetch$0.077 token cost3.11s27,092$0.005 / search$0.001 / fetch
Parallel basicPOST /v1/search mode=basicPOST /v1/extract76.0 ± 0.030s$0.107$0.023 search+fetch$0.084 token cost1.59s32,809$0.005 / search$0.001 / fetch
PerplexityPOST /search search_context_size=highPOST /search search_context_size=high77.7 ± 1.522s$0.083$0.025 search+fetch$0.058 token cost991ms20,062$0.005 / search$0.005 / fetch
Tavily advancedPOST /search search_depth=advancedPOST /extract extract_depth=advanced60.0 ± 2.046s$0.153$0.085 search+fetch$0.069 token cost3.41s26,269$0.016 / search$0.0032 / fetch
Tavily basicPOST /search search_depth=basicPOST /extract extract_depth=basic59.0 ± 1.732s$0.113$0.043 search+fetch$0.070 token cost1.50s27,405$0.008 / search$0.0016 / fetch
TinyFishGET api.search.tinyfish.aiGET api.fetch.tinyfish.ai format=markdown79.0 ± 2.024s$0.050$0.000 search+fetch$0.050 token cost1.32s12,844$0 / search$0 / fetch
YouPOST /v1/search extraction_mode=highlightsPOST /v1/contents55.0 ± 1.727s$0.119$0.027 search+fetch$0.092 token cost638ms42,806$0.005 / search$0.001 / fetch
YouPOST /v1/search extraction_mode=highlights · knowledge=corePOST /v1/contents54.0 ± 2.028s$0.117$0.026 search+fetch$0.091 token cost678ms37,453$0.005 / search$0.001 / fetch

Task completion is mean ± SD of 3 runs; n = 100 tasks. Median task time, median task cost, avg search time, and median task tokens are pooled across those same 3 runs.

[02.b]search only

Perplexity completes most tasks at 77.3%Perplexity is fastest at 18.3 sTinyFish costs least at $0.035

Top 8 per metric
Search-only coding-agent board, listed alphabetically
VendorEndpoint & configurationTask completionMedian task timeMedian task costAvg search timeMedian task tokensAPI list price
BravePOST /res/v1/llm/context38.0 ± 2.620s$0.082$0.025 search$0.057 token cost523ms14,103$0.005 / search
Exa fastPOST /search type=fast66.3 ± 1.520s$0.103$0.035 search$0.068 token cost626ms22,344$0.007 / search
Exa instantPOST /search type=instant61.3 ± 2.921s$0.105$0.035 search$0.070 token cost447ms22,423$0.007 / search
FirecrawlPOST /v2/search70.3 ± 1.528s$0.059$0.025 search$0.034 token cost2.87s7,456$0.005 / search
Linkup fastPOST /v1/search depth=fast43.3 ± 1.127s$0.100$0.025 search$0.075 token cost1.39s24,057$0.005 / search
NimblePOST /v2/search search_depth=lite · full_content=false · focus=general46.7 ± 3.833s$0.043$0.005 search$0.037 token cost3.15s8,713$0.0011 / search
NimblePOST /v2/search search_depth=standard · full_content=false · focus=general42.3 ± 2.525s$0.079$0.025 search$0.054 token cost878ms14,859$0.005 / search
Parallel fastPOST /v1/search mode=fast66.7 ± 1.522s$0.054$0.005 search$0.049 token cost953ms12,460$0.001 / search
Parallel turboPOST /v1/search mode=turbo64.7 ± 2.119s$0.059$0.005 search$0.054 token cost333ms14,130$0.001 / search
PerplexityPOST /search search_context_size=low77.3 ± 2.118s$0.059$0.025 search$0.034 token cost957ms8,765$0.005 / search
Tavily basicPOST /search search_depth=basic51.0 ± 5.326s$0.097$0.040 search$0.057 token cost1.72s16,299$0.008 / search
TinyFishGET api.search.tinyfish.ai59.3 ± 1.528s$0.035$0.000 search$0.035 token cost2.15s7,469$0 / search
YouPOST /v1/search extraction_mode=highlights41.3 ± 3.124s$0.097$0.025 search$0.072 token cost596ms20,333$0.005 / search
YouPOST /v1/search extraction_mode=highlights · knowledge=core38.3 ± 3.221s$0.096$0.025 search$0.071 token cost677ms20,786$0.005 / search

Task completion is mean ± SD of 3 runs; n = 100 tasks. Median task time, median task cost, avg search time, and median task tokens are pooled across those same 3 runs.

Open CodeHarness, scoring and vendor runners in openbenchmarks-labs/web-search-for-coding-agents.github →public dataset30 public tickets on Hugging Face as openbenchmarks/OB-Code-Websearch. Boards are scored on a held-out private set.huggingface →
[02.c] methodology and metric definitions+

How the hard retrieval task is built

Tasks are built from currently live official vendor documentation. Finding those pages is done by a script, not by the coding agent. The search path used here is a neutral third party that is not being evaluated in this benchmark.

  1. Discover. Start from a known list of official docs pages. A script runs ordinary web searches for more pages on the same enterprise SaaS documentation sites. Anything not on those official sites is dropped. No LLM is in this step.
  2. Fetch. A second script downloads the full text of those pages into a local folder. Still no LLM.
  3. Generate. An Authoring Agent reads only those stored pages and proposes product tickets: opaque gold token from the page, capability English that does not name it, and a starter that calls a real but wrong API.
  4. Verify. Automatic keep-gate rejects a ticket if a capable model can solve it without searching, or if the ticket names the ground-truth keywords.
  5. Manual review. A human reads each accepted draft before it is promoted.

The model (gpt-5.6-sol), task, budgets (5 search, 32 turns), and runner remain fixed. When fetch is enabled, web_fetch goes through the same vendor under test as web_search. Each vendor is run three times on both boards.

What each metric means

  • Task completion · mean share of tickets that pass, across three independent runs. A submission passes only if the submitted file compiles, all ground-truth tokens are present, each has an associated # source: URL, and that URL actually appeared in the search or fetch results from that run.
  • Median task time · median end-to-end wall-clock time for the agent to finish the task, pooled across all three repeats.
  • Avg search time · mean latency of each web_search tool call, pooled across all three repeats.
  • Median task tokens · median LLM prompt plus completion tokens per ticket, pooled across all three repeats.
  • API list price· PAYG dollar rate of one search call, and of one fetch call on the search & fetch board.
  • Median task cost · median LLM dollar cost per ticket from Braintrust plus search/fetch API list spend (call counts × unit rate), pooled across all three repeats.

Multi-hop search

Multi-hop search is a question that cannot be answered in one query. Each item combines three or four constraints that are not conclusively present in the model's training data. We tested that in a run with no search tool: the model scored 0. The agent has to plan several searches and combine what they return. Rows are listed alphabetically, with no headline metric or overall performance ranking.

View full multi-hop search benchmark: multi-turn company search →

Web Search only

The agent can issue focused searches and read the returned titles, URLs, and snippets. It cannot fetch page text.

benchmarks/multi-turn/search-onlycomplete

Parallel basic leads F1 at 46.5Brave Search is fastest at 43.5 sTinyFish costs least at $0.212

Top 8 per metric
Search API configurations compared without page fetching
ProviderEndpoint & configurationF1PrecisionRecallExact setMedian timeMean turnsMedian task costAPI list price
Brave SearchGET /res/v1/web/searchbrave28.0 ± 1.766.9 ± 8.519.3 ± 0.50.7 ± 1.343.5 s8.0$0.268$0.070 search$0.200 token cost$0.005 / search
Exa deepPOST /search type=deepexa-deep45.4 ± 2.083.2 ± 3.033.7 ± 1.41.5 ± 1.389.5 s7.4$0.717$0.156 search$0.556 token cost$0.012 / search
Exa instantPOST /search type=instantexa-instant43.3 ± 1.082.6 ± 3.832.2 ± 0.53.0 ± 1.349.8 s7.3$0.652$0.091 search$0.559 token cost$0.007 / search
FirecrawlPOST /v2/searchfirecrawl30.4 ± 1.177.3 ± 4.420.7 ± 0.72.2 ± 0.075.0 s7.9$0.282$0.070 search$0.212 token cost$0.005 / search
Linkup fastPOST /v1/search depth=fastlinkup-fast41.1 ± 1.782.8 ± 1.030.3 ± 1.50.7 ± 1.364.2 s7.4$0.923$0.070 search$0.853 token cost$0.005 / search
Linkup standardPOST /v1/search depth=standardlinkup-standard40.6 ± 0.984.0 ± 7.929.7 ± 1.41.5 ± 2.672.3 s7.3$0.937$0.070 search$0.867 token cost$0.005 / search
NimblePOST /v2/searchsearch_depth=lite · full_content=false · focus=general24.1 ± 1.465.2 ± 2.316.1 ± 1.30.7 ± 1.367.2 s7.9$0.222$0.015 search$0.207 token cost$0.0011 / search
NimblePOST /v2/searchsearch_depth=standard · full_content=false · focus=general30.7 ± 1.470.3 ± 3.321.3 ± 1.10.7 ± 1.353.0 s7.7$0.467$0.070 search$0.398 token cost$0.005 / search
Parallel advancedPOST /v1/search mode=advancedparallel-advanced44.2 ± 1.487.6 ± 3.432.0 ± 1.62.2 ± 0.083.4 s7.3$0.625$0.070 search$0.565 token cost$0.005 / search
Parallel basicPOST /v1/search mode=basicparallel-basic46.5 ± 1.988.7 ± 0.934.4 ± 2.03.7 ± 2.667.6 s7.4$1.120$0.070 search$1.050 token cost$0.005 / search
Parallel fastPOST /v1/search mode=fastparallel-fast38.0 ± 2.079.9 ± 4.627.5 ± 1.41.5 ± 1.353.9 s7.6$0.460$0.014 search$0.447 token cost$0.001 / search
Parallel turboPOST /v1/search mode=turboparallel-turbo34.7 ± 2.480.0 ± 5.224.8 ± 2.41.5 ± 1.346.4 s7.6$0.419$0.014 search$0.405 token cost$0.001 / search
PerplexityPOST /searchsearch_context_size=low37.8 ± 2.179.3 ± 5.726.8 ± 1.42.2 ± 2.248.9 s7.7$0.334$0.070 search$0.264 token cost$0.005 / search
SeltzPOST /v1/search scope=companiesseltz-companies14.5 ± 0.940.0 ± 3.19.4 ± 0.70.0 ± 0.055.2 s7.3$1.751$0.070 search$1.681 token cost$0.005 / search
SERP (RapidAPI)GET google-search74.p.rapidapi.comserp0.4 ± 0.60.7 ± 1.30.3 ± 0.40.0 ± 0.031.4 s7.8$0.103$0.036 search$0.065 token cost$0.003 / search
Tavily advancedPOST /search search_depth=advancedtavily-advanced41.1 ± 2.383.7 ± 4.229.8 ± 1.72.2 ± 0.092.2 s7.4$1.029$0.224 search$0.808 token cost$0.016 / search
Tavily basicPOST /search search_depth=basictavily-basic36.5 ± 3.282.5 ± 1.925.7 ± 2.71.5 ± 1.363.0 s7.7$0.615$0.112 search$0.504 token cost$0.008 / search
TinyFishGET api.search.tinyfish.aitinyfish26.6 ± 1.364.7 ± 1.517.9 ± 0.81.5 ± 1.364.2 s7.9$0.212$0.000 search$0.212 token cost$0 / search
YouPOST /v1/searchextraction_mode=highlights38.6 ± 0.879.3 ± 4.928.3 ± 0.94.4 ± 0.048.2 s7.5$0.971$0.070 search$0.901 token cost$0.005 / search
YouPOST /v1/searchextraction_mode=highlights · knowledge=core38.1 ± 0.680.9 ± 1.627.4 ± 0.93.0 ± 1.347.6 s7.5$0.895$0.070 search$0.827 token cost$0.005 / search

F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor. Median task cost is the median of LLM $ plus search/fetch API $ per agent run. API list price is the PAYG unit rate of the search (and fetch) endpoint the harness calls.

Web Search plus Web Fetch

The same agent can also fetch an exact URL returned by search. Provider-native extraction is used where available.

benchmarks/multi-turn/search-and-fetchcomplete

Exa deep leads F1 at 48.2Brave Search is fastest at 45.1 sTinyFish costs least at $0.219

Top 8 per metric
Search API configurations compared with page fetching enabled
ProviderEndpoint & configurationF1PrecisionRecallExact setMedian timeMean turnsMedian task costAPI list price
Brave SearchGET /res/v1/web/searchbrave29.4 ± 1.673.5 ± 5.820.4 ± 0.91.5 ± 1.345.1 s8.0$0.285$0.060 search+fetch$0.222 token cost$0.005 / search
Exa deepPOST /search type=deepexa-deep48.2 ± 2.189.4 ± 1.236.0 ± 2.22.2 ± 2.295.8 s7.5$0.683$0.156 search+fetch$0.526 token cost$0.012 / search$0.001 / fetch
Exa instantPOST /search type=instantexa-instant44.9 ± 0.985.9 ± 2.533.5 ± 0.95.2 ± 1.352.5 s7.5$0.653$0.091 search+fetch$0.557 token cost$0.007 / search$0.001 / fetch
FirecrawlPOST /v2/searchfirecrawl33.2 ± 2.183.3 ± 0.822.7 ± 1.81.5 ± 1.382.1 s7.9$0.295$0.063 search+fetch$0.230 token cost$0.005 / search$0.0025 / fetch
Linkup fastPOST /v1/search depth=fastlinkup-fast39.9 ± 1.385.3 ± 3.628.6 ± 1.30.7 ± 1.368.0 s7.3$0.903$0.061 search+fetch$0.837 token cost$0.005 / search$0.001 / fetch
Linkup standardPOST /v1/search depth=standardlinkup-standard42.0 ± 1.890.7 ± 2.030.5 ± 2.03.0 ± 3.481.0 s7.5$0.911$0.070 search+fetch$0.848 token cost$0.005 / search$0.001 / fetch
NimblePOST /v2/searchsearch_depth=lite · full_content=false · focus=general · extract formats=[markdown]25.8 ± 2.267.2 ± 10.217.7 ± 1.51.5 ± 1.372.4 s8.0$0.226$0.226 token cost$0.0011 / search
NimblePOST /v2/searchsearch_depth=standard · full_content=false · focus=general · extract formats=[markdown]31.3 ± 1.873.2 ± 1.921.7 ± 1.61.5 ± 1.359.2 s7.9$0.424$0.424 token cost$0.005 / search
Parallel advancedPOST /v1/search mode=advancedparallel-advanced42.2 ± 1.187.6 ± 0.330.1 ± 1.12.2 ± 0.080.9 s7.4$0.599$0.061 search+fetch$0.538 token cost$0.005 / search$0.001 / fetch
Parallel basicPOST /v1/search mode=basicparallel-basic42.3 ± 1.181.3 ± 2.831.3 ± 1.13.0 ± 1.368.6 s7.5$1.089$0.061 search+fetch$1.033 token cost$0.005 / search$0.001 / fetch
Parallel fastPOST /v1/search mode=fastparallel-fast39.3 ± 3.382.3 ± 6.628.2 ± 2.02.2 ± 0.055.1 s7.6$0.441$0.014 search+fetch$0.427 token cost$0.001 / search$0.001 / fetch
Parallel turboPOST /v1/search mode=turboparallel-turbo36.0 ± 3.583.6 ± 4.025.0 ± 2.60.0 ± 0.048.1 s7.7$0.414$0.013 search+fetch$0.400 token cost$0.001 / search$0.001 / fetch
PerplexityPOST /searchsearch_context_size=high46.6 ± 2.087.7 ± 5.934.7 ± 1.12.2 ± 2.253.9 s7.5$0.504$0.070 search+fetch$0.441 token cost$0.005 / search
SeltzPOST /v1/search scope=companiesseltz-companies16.3 ± 1.549.5 ± 2.410.2 ± 1.20.0 ± 0.060.1 s7.5$1.741$0.070 search+fetch$1.671 token cost$0.005 / search
SERP (RapidAPI)GET google-search74.p.rapidapi.comserp0.0 ± 0.00.0 ± 0.00.0 ± 0.00.0 ± 0.033.4 s7.8$0.102$0.039 search+fetch$0.066 token cost$0.003 / search
Tavily advancedPOST /search search_depth=advancedtavily-advanced41.0 ± 1.389.4 ± 6.029.1 ± 0.72.2 ± 0.092.6 s7.3$0.898$0.195 search+fetch$0.685 token cost$0.016 / search$0.0032 / fetch
Tavily basicPOST /search search_depth=basictavily-basic38.1 ± 1.481.2 ± 5.927.8 ± 0.72.2 ± 0.062.0 s7.8$0.610$0.112 search+fetch$0.507 token cost$0.008 / search$0.0032 / fetch
TinyFishGET api.search.tinyfish.aitinyfish30.2 ± 3.570.9 ± 11.320.9 ± 2.50.0 ± 0.063.4 s7.9$0.219$0.000 search+fetch$0.219 token cost$0 / search$0 / fetch
YouPOST /v1/searchextraction_mode=highlights38.5 ± 1.685.2 ± 2.726.9 ± 1.70.7 ± 1.345.6 s7.4$0.864$0.065 search+fetch$0.795 token cost$0.005 / search$0.001 / fetch
YouPOST /v1/searchextraction_mode=highlights · knowledge=core38.5 ± 2.478.2 ± 2.528.3 ± 3.03.7 ± 1.348.5 s7.4$0.894$0.065 search+fetch$0.830 token cost$0.005 / search$0.001 / fetch

F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor. Median task cost is the median of LLM $ plus search/fetch API $ per agent run. API list price is the PAYG unit rate of the search (and fetch) endpoint the harness calls.

open runner + judgeProvider adapters, local run artifacts, and deterministic scoring are public in openbenchmarks-labs/multi-turn-company-search.github →public dataset10 public company-discovery questions and frozen gold labels on Hugging Face as openbenchmarks/OB-Company-Websearch. Board scores use a separate locked 45-question set.huggingface →
[03.a] methodology and metric definitions+

45 company-discovery questions with hand labelled answer sets

Questions were selected from broad intersections in a frozen company census. Twenty-four questions have three constraints and twenty-one have four. The published gold release contains 375 canonical question-company memberships. Reviewers checked companies against every constraint, resolved names and domains to canonical identities, and froze the gold set before scoring.

45 questions

Investor, accelerator, geography, founding-era, and funding constraints.

3 agent runs per question

Independent stochastic agent runs for every provider and question.

2 modes

Search-only and search-plus-fetch are evaluated separately.

5,400 agent runs

Every completed agent run is weighted equally in the reported averages.

One model and one prompt across providers

  • gpt-5.6-sol, medium reasoning effort.
  • Maximum 8 model turns and 14 searches.
  • Maximum two searches per turn and ten results per search.
  • The final response follows one strict JSON schema with company name, domain, cited URLs, and evidence.
  • The agent is instructed not to use prior knowledge as evidence and must complete at least one successful provider search.

Deterministic set comparison

  • Precision · true positives ÷ all returned companies, reported as the three-trial mean ± SD.
  • Recall · true positives ÷ all gold companies, reported as the three-trial mean ± SD.
  • F1 · harmonic mean of precision and recall per agent run, reported as the three-trial mean ± SD.
  • Exact set · share of agent runs with no false positives and no false negatives.

Returned names and domains are resolved to canonical company identities before scoring. The search provider is the comparison variable; the question, prompt, model, budgets, and output schema are held constant.

Best web search API for AI and LLM agents: FAQ

Can the model answer these questions without web search?

No. The model-only baseline is 0. Factual lookup uses latest company facts after the model's cutoff date. Hard retrieval only passes if the patch cites a ground source URL from that run's search results; without a search tool the model cannot provide one. Multi-hop tasks combine three or four constraints that are not conclusively present in training data; that was confirmed in a run with no search tool. A correct answer has to be found by search.

What is the best web search API for AI agents in 2026?

There is no single best web search API for AI agents across tasks. The 2026 roundup ranks the same vendors on three jobs — factual lookup, hard retrieval, and multi-hop search — and keeps the rankings separate.

What is the best web search API for LLM agents in 2026?

Same board: the best web search API for LLM agents depends on the task. Rankings for factual lookup, hard retrieval, and multi-hop search are not averaged. The model is held constant so the comparison isolates the search API.

Exa vs Tavily vs Brave: which web search API is best?

They are measured on the same three tasks, not averaged into one score. Each head-to-head (Exa vs Tavily, Brave vs Exa, Tavily vs Parallel, Linkup vs Firecrawl) shows factual lookup, hard retrieval, and multi-hop.

Is this benchmark independent?

Yes. No vendor pays for inclusion, ranking, or removal. On every task the question set, the agent or extract model, and the judge or scorer are held constant. The only thing that changes is which search API is called.

Full task boards

Best web search API for developers 2026 → · for LLM apps → · for LLM agents →

Best web search APIs for AI agents in 2026: independent benchmarks →

Best search API for coding agents →

Best web search API for AI agents 2026 →

Best web search API for AI agents for fact finding 2026 →

Best web search API for AI agents for deep research 2026 →

Factual lookup: company news →

Hard retrieval: web search for coding agents →

Multi-hop: multi-turn company search →

Factual lookup, ranked for one constraint: most accurate (for LLM apps · for LLM agents) · fastest · cheapest · most token efficient

Hard retrieval, ranked for one constraint: most accurate · fastest · most token efficient

Multi-hop, ranked for one constraint: most accurate · fastest · cheapest

More buying guides: Web search APIs for AI & LLM developers · Web search APIs & MCPs for AI agents and developers · Search tools for AI agents · Search providers for LLM applications · AI search engines for agents · Free web search APIs for AI agents · Web search APIs for RAG · Independent web search API comparison

Use case based guides: Best fast web search API · Best search and scrape API · Best scrape API for AI agents · Best search API for deep research agents · Best search API for company research · Best search API for sales agents · Best search API for coding documentation · Best web search API for grounding · Best search API for news

Head-to-head: Exa vs Tavily · Tavily vs Parallel · Brave Search vs Exa · Linkup vs Firecrawl · Parallel vs Exa · Linkup vs Tavily · Exa vs Perplexity · Brave Search vs Tavily · Exa vs Firecrawl · Brave Search vs Parallel · Perplexity vs Parallel · You vs Parallel · Linkup vs Parallel · Firecrawl vs Parallel · TinyFish vs Parallel · Exa alternatives

Best search APIs: Best Web Search API · Best Search API · Best Web Search API for Agents · Best Search API for Agents · Best Search APIs for AI Agents · Best Search API for AI · Best AI Search API · Best Web Search API for AI Apps · Best Search API for AI Apps · Best Search API for LLM Apps · Best Search API for LLM Agents · Best Search API for RAG · Best Search API for Grounding · Best Search API for AI Agents for Fact-Finding · Best Search API for AI Agents for Deep Research · Best Search API for Developers · Best Search API for AI and LLM Developers · Best Search APIs and MCPs for AI Agents · Best Search API Comparison · Best Semantic Search API

By accuracy, speed, cost and tokens: Most Accurate Web Search API · Most Accurate Web Search API for AI · Most Accurate Search API for AI · Fastest Web Search API for AI · Fastest Web Search API for LLM Agents · Fastest Search API for AI · Fastest Search API for LLM Agents · Cheapest Web Search API for AI · Cheapest Search API for AI · Most Token-Efficient Web Search APIs · Most Token-Efficient Web Search API for AI Agents · Most Token-Efficient Search APIs for AI Agents · Token-Efficient Web Search API for LLM Agents · Token-Efficient Search API for LLM Agents · Best Web Search API by Accuracy · Best Web Search API by Latency · Best Web Search API by Cost · Best Search API by Accuracy · Best Search API by Latency · Best Search API by Cost

Free search APIs: Best Free Web Search API · Best Free Search API · Best Free Search API for AI Agents

Multi-vendor comparisons: Exa vs Tavily vs Brave · Exa vs Tavily vs Brave vs Serper · Exa vs Parallel vs Tavily · Exa vs Tavily vs Parallel vs Linkup · Tavily vs Brave vs Serper · Exa vs Tavily vs Firecrawl · Exa vs Tavily vs Brave vs Perplexity · Exa vs Tavily vs Perplexity

Pricing: Tavily API Pricing · Exa API Pricing · Linkup API Pricing · Parallel Search API Pricing · Brave Search API Pricing

Alternatives: Tavily Alternatives · Brave Search API Alternatives · Parallel Alternatives · Perplexity Alternatives · Linkup Alternatives · Firecrawl Alternatives · You.com Alternatives · Serper Alternatives · SerpApi Alternatives · Bing Search API Alternatives · Google Custom Search API Alternatives

All web search comparisons →

[06] changelog+

Added Nimble lite and standard to Company News, Web Search for Coding Agents, and Multi-turn Company Search. Company News also includes lite with focus=news. Agentic search + fetch uses the separate Extract endpoint; all Search requests use full_content=false.

  • Company news: You highlights scored 90.7%; adding knowledge=core scored 92.0% on 300 questions.
  • Benchmarked Nimble lite and standard plus You highlights with and without knowledge=core on search-only and search + fetch; 100 tasks across three repeats per configuration and mode.
  • Multi-turn company search: both You highlights configurations evaluated on 45 questions across three repeats in both modes, using You Contents for fetch.
  • Replaced Tavily Fast with Tavily Basic on the search-only coding board; evaluated 100 tasks across three repeats.
  • Replaced Tavily ultra-fast with Basic and Advanced on the 300-question company-news benchmark.
  • Factual lookup, coding search, and multi-turn company search tables list providers alphabetically, with search-only and search-plus-fetch results reported separately where applicable.

Re-evaluated TinyFish after updates were rolled out to their GA Fetch endpoint.