45 questions
Investor, accelerator, geography, founding-era, and funding constraints.
Exa (type=fast) leads factual lookup at 99.3% across 300 questions. Exa deep leads hard retrieval at 83.0% on 100 tickets. Parallel basic leads multi-hop search at 46.5% F1 on 45 questions. Independent benchmark of web search APIs for AI agents across 3 tasks. Open source. 300 lookups, 100 hard retrieval, 45 deep research. Last measured 26 Sept 2026.
Why we built this. The best web search API for AI agents and LLM agents depends on the job. A lookup that answers one fact is different from a retrieval use case that has to find answers buried in pages, which is different from a multi-hop search. There is no single best web search API — the winner changes with the tool mode too: Exa deep leads hard retrieval when the agent can fetch pages, Perplexity when it can only read snippets, and on multi-hop it is Parallel basic search-only against Exa deep with fetch. This 2026 board measures the same providers on those three tasks.
Model-only baseline: 0.The model cannot answer any question on this benchmark without web search. That is how the tasks were created: a correct answer has to be found, not recalled. Factual lookup uses latest company facts dated after the model's cutoff. Hard retrieval only passes if the patch cites a ground source URL from that run's search results; without a search tool the model cannot provide one. Multi-hop questions combine three or four constraints that are not conclusively present in training data; we confirmed that in a run with no search tool, where the model still scored 0.
How it is measured.We keep the model and judge constant for all search APIs in each task so the comparison isolates the web search API's performance. Because the model-only score is 0, the ranking is the search API, not the model's memory.
A factual lookup is a single question with one verifiable answer. Search either returns that fact or it does not. Every question is a recent company event dated after the model's cutoff, so the model cannot answer it from memory.
Providers are listed alphabetically. Compare accuracy, answer recall, latency, token usage, and cost side by side, with no headline metric or overall performance ranking.
| Vendor | Endpoint & configuration | Accuracy | AR@1 | AR@5 | $ / 1k correct | Latency | Snippet tokens | Total $ | Official docs | Cost per 1,000 queries |
|---|---|---|---|---|---|---|---|---|---|---|
| Brave Searchweb search | POST /res/v1/llm/contextcount=10 | 94.0% | 81.0% | 94.7% | $5.32 | 601ms | 2,064 | $1.50 | Official docs | $5 / 1kSearch plan · LLM Context |
| Brave Searchweb search | GET /res/v1/web/searchcount=10 · result_filter=web | 93.3% | 79.3% | 91.7% | $5.36 | 630ms | 817 | $1.50 | Official docs | $5 / 1k |
| Exaweb search | POST /searchtype=fast | 99.3% | 95.0% | 99.3% | $7.05 | 652ms | 1,987 | $2.10 | Official docs | $7 / 1ktype=fast · up to 10 results |
| Exaweb search | POST /searchtype=instant | 97.7% | 80.0% | 97.3% | $7.17 | 398ms | 2,128 | $2.10 | Official docs | $7 / 1ktype=instant · up to 10 results |
| Firecrawlweb search | POST /v2/search | 95.3% | 77.7% | 96.7% | $5.24 | 510ms | 678 | $1.50 | Official docs | $5 / 1k2 credits / 10 results |
| Linkupweb search | POST /v1/searchdepth=fast · outputType=searchResults | 96.7% | 79.0% | 94.7% | $5.17 | 1.57s | 3,022 | $1.50 | Official docs | $5 / 1kdepth=fast · searchResults |
| Linkupweb search | POST /v1/searchdepth=standard · outputType=searchResults | 92.0% | 67.3% | 90.3% | $5.43 | 2.55s | 2,983 | $1.50 | Official docs | $5 / 1kdepth=standard · searchResults |
| Nimbleweb search | POST /v2/searchsearch_depth=lite · full_content=false · focus=general | 75.3% | 64.0% | 75.3% | $1.46 | 3.15s | 413 | $0.33 | Official docs | $1.1 / 1ksearch_depth=lite · full_content=false |
| Nimbleweb search | POST /v2/searchsearch_depth=standard · full_content=false · focus=general | 93.0% | 86.3% | 94.3% | $5.38 | 861ms | 2,730 | $1.50 | Official docs | $5 / 1ksearch_depth=standard · full_content=false |
| Parallelweb search | POST /v1/searchmode=basic | 93.3% | 55.0% | 92.0% | $5.36 | 1.68s | 2,330 | $1.50 | Official docs | $5 / 1kmode=basic · 10 results |
| Parallelweb search | POST /v1/searchmode=fast | 86.0% | 44.3% | 79.0% | $1.16 | 942ms | 1,839 | $0.30 | Official docs | $1 / 1kmode=fast · 10 results |
| Parallelweb search | POST /v1/searchmode=turbo | 71.3% | 45.3% | 66.0% | $1.40 | 348ms | 1,853 | $0.30 | Official docs | $1 / 1kmode=turbo · 10 results |
| Perplexityweb search | POST /searchsearch_context_size=low | 97.3% | 91.7% | 98.3% | $5.14 | 1.38s | 476 | $1.50 | Official docs | $5 / 1kSearch API · POST /search |
| SERPgoogle search | GET google-search74.p.rapidapi.comlimit=10 | 96.0% | 78.0% | 95.0% | $3.13 | 751ms | 497 | $0.90 | Official docs | $3 / 1kPro overage $0.003/request |
| Tavilyweb search | POST /searchsearch_depth=advanced | 93.0% | 65.0% | 92.7% | $17.20 | 4.29s | 2,210 | $4.80 | Official docs | $16 / 1k2 credits · $0.008 PAYG |
| Tavilyweb search | POST /searchsearch_depth=basic | 87.7% | 82.0% | 87.7% | $9.13 | 1.88s | 1,639 | $2.40 | Official docs | $8 / 1k1 credit · $0.008 PAYG |
| TinyFishweb search | GET api.search.tinyfish.aifree · 30 req/min cap | 92.0% | 74.3% | 90.7% | Free30 req/min cap | 2.62s | 441 | Free | Official docs | $0free · 30 req/min, 500/hour |
| Youweb search | POST /v1/searchextraction_mode=highlights | 90.7% | 72.0% | 89.7% | $5.51 | 628ms | 2,837 | $1.50 | Official docs | $5 / 1khighlights included in Web Search |
| Youweb search | POST /v1/searchextraction_mode=highlights · knowledge=core | 92.0% | 67.3% | 89.7% | $5.43 | 889ms | 2,862 | $1.50 | Official docs | $5 / 1khighlights · knowledge=core · same Web Search price assumption |
$ / 1k correct is $ / 1k queries divided by extracted-answer accuracy. Providers are listed alphabetically; no single metric determines their order. Cost per 1,000 queries is the published PAYG list price, linked to the vendor pricing page. Not promotional packs or volume discounts. TinyFish Search is free with a 30 requests/min cap — $0 is not unlimited throughput.
We measure the web search API. Every endpoint answers the same 300 lookup questions. The query, the model that reads the results and extracts an answer (gpt-5.6-terra, medium), and the judge (claude-opus-5 via Amazon Bedrock) are held constant. The only thing that changes is which search API is called.
gpt-5.6-terra, medium) reads title + excerpt only and extracts an answer. A separate post-hoc judge from an independent model family (claude-opus-5 via Amazon Bedrock) scores accuracy and AR@K against the human labelled ground truth. It never fetches the live page.A 100-question public set is on Hugging Face as openbenchmarks/OB-News-Websearch. Board scores use the locked 300-question set. Harness, scoring and vendor runners are in openbenchmarks-labs/factual-lookup-company-news-search.
Accuracy · gpt-5.6-terra (medium) writes an answer from the returned titles and snippets. The judge scores that written answer against human labelled ground truth.AR@K · answer recall at K. Whether any of the top K vendor snippets already contained the human labelled ground truth. No extract.AR@1 · the first snippet already contained the human labelled ground truth.AR@5 · whether any one of the top 5 snippets already contained the full human labelled ground truth.Latency · mean wall time of the search request, not the extract or score calls.Total $ · search list price for the run. Judge tokens are not in this column.Cost per 1,000 queries · published PAYG list price of the endpoint we called, linked to the vendor pricing page. Per-request search APIs are $ / 1,000 queries. Autobound and Datahyena bill credits, so that cell is $ / 1,000 credits.$ / 1k correct · $ / 1k queries divided by extracted-answer accuracy. Same cost-to-quality figure as the cheapest page.Hard retrieval is finding an answer that is seated deep within a unique page, across very similar pages. The snippet is not enough, and the model cannot guess the detail. Both boards: search & fetch, then search-only.
View full hard retrieval benchmark: web search for coding agents →
| Vendor | Endpoint & configuration | Task completion | Median task time | Median task cost | Avg search time | Median task tokens | API list price |
|---|---|---|---|---|---|---|---|
| Exa auto | POST /search type=autoPOST /contents | 81.7 ± 1.1 | 23s | $0.107$0.032 search+fetch$0.074 token cost | 1.19s | 27,433 | $0.007 / search$0.001 / fetch |
| Exa deep | POST /search type=deepPOST /contents | 83.0 ± 1.0 | 37s | $0.127$0.057 search+fetch$0.070 token cost | 3.97s | 23,660 | $0.012 / search$0.001 / fetch |
| Firecrawl | POST /v2/searchPOST /v2/scrape | 76.0 ± 1.0 | 34s | $0.082$0.028 search+fetch$0.055 token cost | 2.81s | 17,379 | $0.005 / search$0.0025 / fetch |
| Linkup standard | POST /v1/search depth=standardPOST /v1/fetch mode=standard | 48.3 ± 4.9 | 48s | $0.147$0.038 search+fetch$0.108 token cost | 2.00s | 57,791 | $0.005 / search$0.005 / fetch |
| Nimble | POST /v2/search search_depth=lite · full_content=false · focus=generalPOST /v2/extract formats=[markdown] | 60.3 ± 2.5 | 40s | $0.054$0.054 token cost | 1.93s | 16,828 | $0.0011 / search |
| Nimble | POST /v2/search search_depth=standard · full_content=false · focus=generalPOST /v2/extract formats=[markdown] | 45.0 ± 1.0 | 44s | $0.078$0.078 token cost | 905ms | 28,586 | $0.005 / search |
| Parallel advanced | POST /v1/search mode=advancedPOST /v1/extract | 77.0 ± 1.0 | 33s | $0.096$0.019 search+fetch$0.077 token cost | 3.11s | 27,092 | $0.005 / search$0.001 / fetch |
| Parallel basic | POST /v1/search mode=basicPOST /v1/extract | 76.0 ± 0.0 | 30s | $0.107$0.023 search+fetch$0.084 token cost | 1.59s | 32,809 | $0.005 / search$0.001 / fetch |
| Perplexity | POST /search search_context_size=highPOST /search search_context_size=high | 77.7 ± 1.5 | 22s | $0.083$0.025 search+fetch$0.058 token cost | 991ms | 20,062 | $0.005 / search$0.005 / fetch |
| Tavily advanced | POST /search search_depth=advancedPOST /extract extract_depth=advanced | 60.0 ± 2.0 | 46s | $0.153$0.085 search+fetch$0.069 token cost | 3.41s | 26,269 | $0.016 / search$0.0032 / fetch |
| Tavily basic | POST /search search_depth=basicPOST /extract extract_depth=basic | 59.0 ± 1.7 | 32s | $0.113$0.043 search+fetch$0.070 token cost | 1.50s | 27,405 | $0.008 / search$0.0016 / fetch |
| TinyFish | GET api.search.tinyfish.aiGET api.fetch.tinyfish.ai format=markdown | 79.0 ± 2.0 | 24s | $0.050$0.000 search+fetch$0.050 token cost | 1.32s | 12,844 | $0 / search$0 / fetch |
| You | POST /v1/search extraction_mode=highlightsPOST /v1/contents | 55.0 ± 1.7 | 27s | $0.119$0.027 search+fetch$0.092 token cost | 638ms | 42,806 | $0.005 / search$0.001 / fetch |
| You | POST /v1/search extraction_mode=highlights · knowledge=corePOST /v1/contents | 54.0 ± 2.0 | 28s | $0.117$0.026 search+fetch$0.091 token cost | 678ms | 37,453 | $0.005 / search$0.001 / fetch |
Task completion is mean ± SD of 3 runs; n = 100 tasks. Median task time, median task cost, avg search time, and median task tokens are pooled across those same 3 runs.
| Vendor | Endpoint & configuration | Task completion | Median task time | Median task cost | Avg search time | Median task tokens | API list price |
|---|---|---|---|---|---|---|---|
| Brave | POST /res/v1/llm/context | 38.0 ± 2.6 | 20s | $0.082$0.025 search$0.057 token cost | 523ms | 14,103 | $0.005 / search |
| Exa fast | POST /search type=fast | 66.3 ± 1.5 | 20s | $0.103$0.035 search$0.068 token cost | 626ms | 22,344 | $0.007 / search |
| Exa instant | POST /search type=instant | 61.3 ± 2.9 | 21s | $0.105$0.035 search$0.070 token cost | 447ms | 22,423 | $0.007 / search |
| Firecrawl | POST /v2/search | 70.3 ± 1.5 | 28s | $0.059$0.025 search$0.034 token cost | 2.87s | 7,456 | $0.005 / search |
| Linkup fast | POST /v1/search depth=fast | 43.3 ± 1.1 | 27s | $0.100$0.025 search$0.075 token cost | 1.39s | 24,057 | $0.005 / search |
| Nimble | POST /v2/search search_depth=lite · full_content=false · focus=general | 46.7 ± 3.8 | 33s | $0.043$0.005 search$0.037 token cost | 3.15s | 8,713 | $0.0011 / search |
| Nimble | POST /v2/search search_depth=standard · full_content=false · focus=general | 42.3 ± 2.5 | 25s | $0.079$0.025 search$0.054 token cost | 878ms | 14,859 | $0.005 / search |
| Parallel fast | POST /v1/search mode=fast | 66.7 ± 1.5 | 22s | $0.054$0.005 search$0.049 token cost | 953ms | 12,460 | $0.001 / search |
| Parallel turbo | POST /v1/search mode=turbo | 64.7 ± 2.1 | 19s | $0.059$0.005 search$0.054 token cost | 333ms | 14,130 | $0.001 / search |
| Perplexity | POST /search search_context_size=low | 77.3 ± 2.1 | 18s | $0.059$0.025 search$0.034 token cost | 957ms | 8,765 | $0.005 / search |
| Tavily basic | POST /search search_depth=basic | 51.0 ± 5.3 | 26s | $0.097$0.040 search$0.057 token cost | 1.72s | 16,299 | $0.008 / search |
| TinyFish | GET api.search.tinyfish.ai | 59.3 ± 1.5 | 28s | $0.035$0.000 search$0.035 token cost | 2.15s | 7,469 | $0 / search |
| You | POST /v1/search extraction_mode=highlights | 41.3 ± 3.1 | 24s | $0.097$0.025 search$0.072 token cost | 596ms | 20,333 | $0.005 / search |
| You | POST /v1/search extraction_mode=highlights · knowledge=core | 38.3 ± 3.2 | 21s | $0.096$0.025 search$0.071 token cost | 677ms | 20,786 | $0.005 / search |
Task completion is mean ± SD of 3 runs; n = 100 tasks. Median task time, median task cost, avg search time, and median task tokens are pooled across those same 3 runs.
Tasks are built from currently live official vendor documentation. Finding those pages is done by a script, not by the coding agent. The search path used here is a neutral third party that is not being evaluated in this benchmark.
The model (gpt-5.6-sol), task, budgets (5 search, 32 turns), and runner remain fixed. When fetch is enabled, web_fetch goes through the same vendor under test as web_search. Each vendor is run three times on both boards.
Task completion · mean share of tickets that pass, across three independent runs. A submission passes only if the submitted file compiles, all ground-truth tokens are present, each has an associated # source: URL, and that URL actually appeared in the search or fetch results from that run.Median task time · median end-to-end wall-clock time for the agent to finish the task, pooled across all three repeats.Avg search time · mean latency of each web_search tool call, pooled across all three repeats.Median task tokens · median LLM prompt plus completion tokens per ticket, pooled across all three repeats.API list price· PAYG dollar rate of one search call, and of one fetch call on the search & fetch board.Median task cost · median LLM dollar cost per ticket from Braintrust plus search/fetch API list spend (call counts × unit rate), pooled across all three repeats.Multi-hop search is a question that cannot be answered in one query. Each item combines three or four constraints that are not conclusively present in the model's training data. We tested that in a run with no search tool: the model scored 0. The agent has to plan several searches and combine what they return. Rows are listed alphabetically, with no headline metric or overall performance ranking.
View full multi-hop search benchmark: multi-turn company search →
The agent can issue focused searches and read the returned titles, URLs, and snippets. It cannot fetch page text.
| Provider | Endpoint & configuration | F1 | Precision | Recall | Exact set | Median time | Mean turns | Median task cost | API list price |
|---|---|---|---|---|---|---|---|---|---|
| Brave Search | GET /res/v1/web/searchbrave | 28.0 ± 1.7 | 66.9 ± 8.5 | 19.3 ± 0.5 | 0.7 ± 1.3 | 43.5 s | 8.0 | $0.268$0.070 search$0.200 token cost | $0.005 / search |
| Exa deep | POST /search type=deepexa-deep | 45.4 ± 2.0 | 83.2 ± 3.0 | 33.7 ± 1.4 | 1.5 ± 1.3 | 89.5 s | 7.4 | $0.717$0.156 search$0.556 token cost | $0.012 / search |
| Exa instant | POST /search type=instantexa-instant | 43.3 ± 1.0 | 82.6 ± 3.8 | 32.2 ± 0.5 | 3.0 ± 1.3 | 49.8 s | 7.3 | $0.652$0.091 search$0.559 token cost | $0.007 / search |
| Firecrawl | POST /v2/searchfirecrawl | 30.4 ± 1.1 | 77.3 ± 4.4 | 20.7 ± 0.7 | 2.2 ± 0.0 | 75.0 s | 7.9 | $0.282$0.070 search$0.212 token cost | $0.005 / search |
| Linkup fast | POST /v1/search depth=fastlinkup-fast | 41.1 ± 1.7 | 82.8 ± 1.0 | 30.3 ± 1.5 | 0.7 ± 1.3 | 64.2 s | 7.4 | $0.923$0.070 search$0.853 token cost | $0.005 / search |
| Linkup standard | POST /v1/search depth=standardlinkup-standard | 40.6 ± 0.9 | 84.0 ± 7.9 | 29.7 ± 1.4 | 1.5 ± 2.6 | 72.3 s | 7.3 | $0.937$0.070 search$0.867 token cost | $0.005 / search |
| Nimble | POST /v2/searchsearch_depth=lite · full_content=false · focus=general | 24.1 ± 1.4 | 65.2 ± 2.3 | 16.1 ± 1.3 | 0.7 ± 1.3 | 67.2 s | 7.9 | $0.222$0.015 search$0.207 token cost | $0.0011 / search |
| Nimble | POST /v2/searchsearch_depth=standard · full_content=false · focus=general | 30.7 ± 1.4 | 70.3 ± 3.3 | 21.3 ± 1.1 | 0.7 ± 1.3 | 53.0 s | 7.7 | $0.467$0.070 search$0.398 token cost | $0.005 / search |
| Parallel advanced | POST /v1/search mode=advancedparallel-advanced | 44.2 ± 1.4 | 87.6 ± 3.4 | 32.0 ± 1.6 | 2.2 ± 0.0 | 83.4 s | 7.3 | $0.625$0.070 search$0.565 token cost | $0.005 / search |
| Parallel basic | POST /v1/search mode=basicparallel-basic | 46.5 ± 1.9 | 88.7 ± 0.9 | 34.4 ± 2.0 | 3.7 ± 2.6 | 67.6 s | 7.4 | $1.120$0.070 search$1.050 token cost | $0.005 / search |
| Parallel fast | POST /v1/search mode=fastparallel-fast | 38.0 ± 2.0 | 79.9 ± 4.6 | 27.5 ± 1.4 | 1.5 ± 1.3 | 53.9 s | 7.6 | $0.460$0.014 search$0.447 token cost | $0.001 / search |
| Parallel turbo | POST /v1/search mode=turboparallel-turbo | 34.7 ± 2.4 | 80.0 ± 5.2 | 24.8 ± 2.4 | 1.5 ± 1.3 | 46.4 s | 7.6 | $0.419$0.014 search$0.405 token cost | $0.001 / search |
| Perplexity | POST /searchsearch_context_size=low | 37.8 ± 2.1 | 79.3 ± 5.7 | 26.8 ± 1.4 | 2.2 ± 2.2 | 48.9 s | 7.7 | $0.334$0.070 search$0.264 token cost | $0.005 / search |
| Seltz | POST /v1/search scope=companiesseltz-companies | 14.5 ± 0.9 | 40.0 ± 3.1 | 9.4 ± 0.7 | 0.0 ± 0.0 | 55.2 s | 7.3 | $1.751$0.070 search$1.681 token cost | $0.005 / search |
| SERP (RapidAPI) | GET google-search74.p.rapidapi.comserp | 0.4 ± 0.6 | 0.7 ± 1.3 | 0.3 ± 0.4 | 0.0 ± 0.0 | 31.4 s | 7.8 | $0.103$0.036 search$0.065 token cost | $0.003 / search |
| Tavily advanced | POST /search search_depth=advancedtavily-advanced | 41.1 ± 2.3 | 83.7 ± 4.2 | 29.8 ± 1.7 | 2.2 ± 0.0 | 92.2 s | 7.4 | $1.029$0.224 search$0.808 token cost | $0.016 / search |
| Tavily basic | POST /search search_depth=basictavily-basic | 36.5 ± 3.2 | 82.5 ± 1.9 | 25.7 ± 2.7 | 1.5 ± 1.3 | 63.0 s | 7.7 | $0.615$0.112 search$0.504 token cost | $0.008 / search |
| TinyFish | GET api.search.tinyfish.aitinyfish | 26.6 ± 1.3 | 64.7 ± 1.5 | 17.9 ± 0.8 | 1.5 ± 1.3 | 64.2 s | 7.9 | $0.212$0.000 search$0.212 token cost | $0 / search |
| You | POST /v1/searchextraction_mode=highlights | 38.6 ± 0.8 | 79.3 ± 4.9 | 28.3 ± 0.9 | 4.4 ± 0.0 | 48.2 s | 7.5 | $0.971$0.070 search$0.901 token cost | $0.005 / search |
| You | POST /v1/searchextraction_mode=highlights · knowledge=core | 38.1 ± 0.6 | 80.9 ± 1.6 | 27.4 ± 0.9 | 3.0 ± 1.3 | 47.6 s | 7.5 | $0.895$0.070 search$0.827 token cost | $0.005 / search |
F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor. Median task cost is the median of LLM $ plus search/fetch API $ per agent run. API list price is the PAYG unit rate of the search (and fetch) endpoint the harness calls.
The same agent can also fetch an exact URL returned by search. Provider-native extraction is used where available.
| Provider | Endpoint & configuration | F1 | Precision | Recall | Exact set | Median time | Mean turns | Median task cost | API list price |
|---|---|---|---|---|---|---|---|---|---|
| Brave Search | GET /res/v1/web/searchbrave | 29.4 ± 1.6 | 73.5 ± 5.8 | 20.4 ± 0.9 | 1.5 ± 1.3 | 45.1 s | 8.0 | $0.285$0.060 search+fetch$0.222 token cost | $0.005 / search |
| Exa deep | POST /search type=deepexa-deep | 48.2 ± 2.1 | 89.4 ± 1.2 | 36.0 ± 2.2 | 2.2 ± 2.2 | 95.8 s | 7.5 | $0.683$0.156 search+fetch$0.526 token cost | $0.012 / search$0.001 / fetch |
| Exa instant | POST /search type=instantexa-instant | 44.9 ± 0.9 | 85.9 ± 2.5 | 33.5 ± 0.9 | 5.2 ± 1.3 | 52.5 s | 7.5 | $0.653$0.091 search+fetch$0.557 token cost | $0.007 / search$0.001 / fetch |
| Firecrawl | POST /v2/searchfirecrawl | 33.2 ± 2.1 | 83.3 ± 0.8 | 22.7 ± 1.8 | 1.5 ± 1.3 | 82.1 s | 7.9 | $0.295$0.063 search+fetch$0.230 token cost | $0.005 / search$0.0025 / fetch |
| Linkup fast | POST /v1/search depth=fastlinkup-fast | 39.9 ± 1.3 | 85.3 ± 3.6 | 28.6 ± 1.3 | 0.7 ± 1.3 | 68.0 s | 7.3 | $0.903$0.061 search+fetch$0.837 token cost | $0.005 / search$0.001 / fetch |
| Linkup standard | POST /v1/search depth=standardlinkup-standard | 42.0 ± 1.8 | 90.7 ± 2.0 | 30.5 ± 2.0 | 3.0 ± 3.4 | 81.0 s | 7.5 | $0.911$0.070 search+fetch$0.848 token cost | $0.005 / search$0.001 / fetch |
| Nimble | POST /v2/searchsearch_depth=lite · full_content=false · focus=general · extract formats=[markdown] | 25.8 ± 2.2 | 67.2 ± 10.2 | 17.7 ± 1.5 | 1.5 ± 1.3 | 72.4 s | 8.0 | $0.226$0.226 token cost | $0.0011 / search |
| Nimble | POST /v2/searchsearch_depth=standard · full_content=false · focus=general · extract formats=[markdown] | 31.3 ± 1.8 | 73.2 ± 1.9 | 21.7 ± 1.6 | 1.5 ± 1.3 | 59.2 s | 7.9 | $0.424$0.424 token cost | $0.005 / search |
| Parallel advanced | POST /v1/search mode=advancedparallel-advanced | 42.2 ± 1.1 | 87.6 ± 0.3 | 30.1 ± 1.1 | 2.2 ± 0.0 | 80.9 s | 7.4 | $0.599$0.061 search+fetch$0.538 token cost | $0.005 / search$0.001 / fetch |
| Parallel basic | POST /v1/search mode=basicparallel-basic | 42.3 ± 1.1 | 81.3 ± 2.8 | 31.3 ± 1.1 | 3.0 ± 1.3 | 68.6 s | 7.5 | $1.089$0.061 search+fetch$1.033 token cost | $0.005 / search$0.001 / fetch |
| Parallel fast | POST /v1/search mode=fastparallel-fast | 39.3 ± 3.3 | 82.3 ± 6.6 | 28.2 ± 2.0 | 2.2 ± 0.0 | 55.1 s | 7.6 | $0.441$0.014 search+fetch$0.427 token cost | $0.001 / search$0.001 / fetch |
| Parallel turbo | POST /v1/search mode=turboparallel-turbo | 36.0 ± 3.5 | 83.6 ± 4.0 | 25.0 ± 2.6 | 0.0 ± 0.0 | 48.1 s | 7.7 | $0.414$0.013 search+fetch$0.400 token cost | $0.001 / search$0.001 / fetch |
| Perplexity | POST /searchsearch_context_size=high | 46.6 ± 2.0 | 87.7 ± 5.9 | 34.7 ± 1.1 | 2.2 ± 2.2 | 53.9 s | 7.5 | $0.504$0.070 search+fetch$0.441 token cost | $0.005 / search |
| Seltz | POST /v1/search scope=companiesseltz-companies | 16.3 ± 1.5 | 49.5 ± 2.4 | 10.2 ± 1.2 | 0.0 ± 0.0 | 60.1 s | 7.5 | $1.741$0.070 search+fetch$1.671 token cost | $0.005 / search |
| SERP (RapidAPI) | GET google-search74.p.rapidapi.comserp | 0.0 ± 0.0 | 0.0 ± 0.0 | 0.0 ± 0.0 | 0.0 ± 0.0 | 33.4 s | 7.8 | $0.102$0.039 search+fetch$0.066 token cost | $0.003 / search |
| Tavily advanced | POST /search search_depth=advancedtavily-advanced | 41.0 ± 1.3 | 89.4 ± 6.0 | 29.1 ± 0.7 | 2.2 ± 0.0 | 92.6 s | 7.3 | $0.898$0.195 search+fetch$0.685 token cost | $0.016 / search$0.0032 / fetch |
| Tavily basic | POST /search search_depth=basictavily-basic | 38.1 ± 1.4 | 81.2 ± 5.9 | 27.8 ± 0.7 | 2.2 ± 0.0 | 62.0 s | 7.8 | $0.610$0.112 search+fetch$0.507 token cost | $0.008 / search$0.0032 / fetch |
| TinyFish | GET api.search.tinyfish.aitinyfish | 30.2 ± 3.5 | 70.9 ± 11.3 | 20.9 ± 2.5 | 0.0 ± 0.0 | 63.4 s | 7.9 | $0.219$0.000 search+fetch$0.219 token cost | $0 / search$0 / fetch |
| You | POST /v1/searchextraction_mode=highlights | 38.5 ± 1.6 | 85.2 ± 2.7 | 26.9 ± 1.7 | 0.7 ± 1.3 | 45.6 s | 7.4 | $0.864$0.065 search+fetch$0.795 token cost | $0.005 / search$0.001 / fetch |
| You | POST /v1/searchextraction_mode=highlights · knowledge=core | 38.5 ± 2.4 | 78.2 ± 2.5 | 28.3 ± 3.0 | 3.7 ± 1.3 | 48.5 s | 7.4 | $0.894$0.065 search+fetch$0.830 token cost | $0.005 / search$0.001 / fetch |
F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor. Median task cost is the median of LLM $ plus search/fetch API $ per agent run. API list price is the PAYG unit rate of the search (and fetch) endpoint the harness calls.
Questions were selected from broad intersections in a frozen company census. Twenty-four questions have three constraints and twenty-one have four. The published gold release contains 375 canonical question-company memberships. Reviewers checked companies against every constraint, resolved names and domains to canonical identities, and froze the gold set before scoring.
Investor, accelerator, geography, founding-era, and funding constraints.
Independent stochastic agent runs for every provider and question.
Search-only and search-plus-fetch are evaluated separately.
Every completed agent run is weighted equally in the reported averages.
Precision · true positives ÷ all returned companies, reported as the three-trial mean ± SD.Recall · true positives ÷ all gold companies, reported as the three-trial mean ± SD.F1 · harmonic mean of precision and recall per agent run, reported as the three-trial mean ± SD.Exact set · share of agent runs with no false positives and no false negatives.Returned names and domains are resolved to canonical company identities before scoring. The search provider is the comparison variable; the question, prompt, model, budgets, and output schema are held constant.
No. The model-only baseline is 0. Factual lookup uses latest company facts after the model's cutoff date. Hard retrieval only passes if the patch cites a ground source URL from that run's search results; without a search tool the model cannot provide one. Multi-hop tasks combine three or four constraints that are not conclusively present in training data; that was confirmed in a run with no search tool. A correct answer has to be found by search.
There is no single best web search API for AI agents across tasks. The 2026 roundup ranks the same vendors on three jobs — factual lookup, hard retrieval, and multi-hop search — and keeps the rankings separate.
Same board: the best web search API for LLM agents depends on the task. Rankings for factual lookup, hard retrieval, and multi-hop search are not averaged. The model is held constant so the comparison isolates the search API.
They are measured on the same three tasks, not averaged into one score. Each head-to-head (Exa vs Tavily, Brave vs Exa, Tavily vs Parallel, Linkup vs Firecrawl) shows factual lookup, hard retrieval, and multi-hop.
Yes. No vendor pays for inclusion, ranking, or removal. On every task the question set, the agent or extract model, and the judge or scorer are held constant. The only thing that changes is which search API is called.
Best web search API for developers 2026 → · for LLM apps → · for LLM agents →
Best web search APIs for AI agents in 2026: independent benchmarks →
Best search API for coding agents →
Best web search API for AI agents 2026 →
Best web search API for AI agents for fact finding 2026 →
Best web search API for AI agents for deep research 2026 →
Factual lookup: company news →
Hard retrieval: web search for coding agents →
Multi-hop: multi-turn company search →
Factual lookup, ranked for one constraint: most accurate (for LLM apps · for LLM agents) · fastest · cheapest · most token efficient
Hard retrieval, ranked for one constraint: most accurate · fastest · most token efficient
Multi-hop, ranked for one constraint: most accurate · fastest · cheapest
More buying guides: Web search APIs for AI & LLM developers · Web search APIs & MCPs for AI agents and developers · Search tools for AI agents · Search providers for LLM applications · AI search engines for agents · Free web search APIs for AI agents · Web search APIs for RAG · Independent web search API comparison
Use case based guides: Best fast web search API · Best search and scrape API · Best scrape API for AI agents · Best search API for deep research agents · Best search API for company research · Best search API for sales agents · Best search API for coding documentation · Best web search API for grounding · Best search API for news
Head-to-head: Exa vs Tavily · Tavily vs Parallel · Brave Search vs Exa · Linkup vs Firecrawl · Parallel vs Exa · Linkup vs Tavily · Exa vs Perplexity · Brave Search vs Tavily · Exa vs Firecrawl · Brave Search vs Parallel · Perplexity vs Parallel · You vs Parallel · Linkup vs Parallel · Firecrawl vs Parallel · TinyFish vs Parallel · Exa alternatives
Best search APIs: Best Web Search API · Best Search API · Best Web Search API for Agents · Best Search API for Agents · Best Search APIs for AI Agents · Best Search API for AI · Best AI Search API · Best Web Search API for AI Apps · Best Search API for AI Apps · Best Search API for LLM Apps · Best Search API for LLM Agents · Best Search API for RAG · Best Search API for Grounding · Best Search API for AI Agents for Fact-Finding · Best Search API for AI Agents for Deep Research · Best Search API for Developers · Best Search API for AI and LLM Developers · Best Search APIs and MCPs for AI Agents · Best Search API Comparison · Best Semantic Search API
By accuracy, speed, cost and tokens: Most Accurate Web Search API · Most Accurate Web Search API for AI · Most Accurate Search API for AI · Fastest Web Search API for AI · Fastest Web Search API for LLM Agents · Fastest Search API for AI · Fastest Search API for LLM Agents · Cheapest Web Search API for AI · Cheapest Search API for AI · Most Token-Efficient Web Search APIs · Most Token-Efficient Web Search API for AI Agents · Most Token-Efficient Search APIs for AI Agents · Token-Efficient Web Search API for LLM Agents · Token-Efficient Search API for LLM Agents · Best Web Search API by Accuracy · Best Web Search API by Latency · Best Web Search API by Cost · Best Search API by Accuracy · Best Search API by Latency · Best Search API by Cost
Free search APIs: Best Free Web Search API · Best Free Search API · Best Free Search API for AI Agents
Multi-vendor comparisons: Exa vs Tavily vs Brave · Exa vs Tavily vs Brave vs Serper · Exa vs Parallel vs Tavily · Exa vs Tavily vs Parallel vs Linkup · Tavily vs Brave vs Serper · Exa vs Tavily vs Firecrawl · Exa vs Tavily vs Brave vs Perplexity · Exa vs Tavily vs Perplexity
Pricing: Tavily API Pricing · Exa API Pricing · Linkup API Pricing · Parallel Search API Pricing · Brave Search API Pricing
Alternatives: Tavily Alternatives · Brave Search API Alternatives · Parallel Alternatives · Perplexity Alternatives · Linkup Alternatives · Firecrawl Alternatives · You.com Alternatives · Serper Alternatives · SerpApi Alternatives · Bing Search API Alternatives · Google Custom Search API Alternatives
Added Nimble lite and standard to Company News, Web Search for Coding Agents, and Multi-turn Company Search. Company News also includes lite with focus=news. Agentic search + fetch uses the separate Extract endpoint; all Search requests use full_content=false.
Re-evaluated TinyFish after updates were rolled out to their GA Fetch endpoint.