benchmarks/web-search/web search for coding agents
100 tasks · search & fetch · search-only · last updated 26 Sept 2026

Web Search for Coding Agents Benchmark

Exa deep leads with fetch at 83.0% task completion on 100 tickets. Perplexity leads search-only at 77.3%. Different APIs lead depending on whether the agent can open pages, so the two boards are ranked separately and never averaged.

Why we built this. This is the hard retrieval task of the web search benchmark. We evaluate web search APIs for coding agents and developers using realistic product tickets. We test search APIs on hard enterprise SaaS documentation. The agent receives a small existing httpxclient and must patch it against the vendor's current official API documentation. The ticket describes the desired capability in natural language, but deliberately does not reveal the exact implementation details.

Model-only baseline: 0. A submission only passes if the patch compiles, the ground-truth tokens are present, and each is grounded in a # source:URL that actually appeared in that run's search or fetch results. Without a search tool the model cannot provide that URL, so it cannot pass. The ticket also does not name the implementation detail, so the path cannot be recalled.

Sample task. A sales team needs to price a quote in ServiceNow Sales CRM. The file the agent gets already makes a live ServiceNow request, but it is the default help-desk call: read the incident table. The gold path is one endpoint.

ticket

Price this quote using the current Sales CRM pricing engine. The starter still reads the incident table, which is the help-desk API for listing support tickets — that cannot compute a price. Look up the official docs and switch the call.

what the docs actually require

POST /api/sn_csm_pricing/v1/pricingengine/computePrice

The inherited code would run today: GET /api/now/table/incident. That returns help-desk tickets. It does not price a quote. The path above lives in Sales CRM pricing docs and is not written anywhere in the ticket.

starter → required patch
@@ price.pydef run() -> dict:-     response = httpx.get(-         f"{BASE}/api/now/table/incident",+     response = httpx.post(+         # source: https://www.servicenow.com/docs/.../sales_crm_pricing-POST-compute-price.html+         f"{BASE}/api/sn_csm_pricing/v1/pricingengine/computePrice",        headers={            "Authorization": f"Bearer {TOKEN}",            "Accept": "application/json",        },

Red is the inherited client. Green is the change after the agent finds the official docs. The # source:comment has to be a docs URL that actually showed up in that run's search or fetch results and not a URL hallucinated by the model.

How it is measured. The model receives only the ticket and the starter code. The search or fetch provider is varied while the model (gpt-5.6-sol), task, budgets (5 search, 32 turns), and runner remain fixed. A submission passes only if the submitted file compiles, all ground-truth tokens are present, each has an associated # source: URL, and that URL actually appeared in the search or fetch results from that run. Both boards are published: search & fetch, then search-only.

Vendors benchmarked. Exa deep, Exa auto, TinyFish, Perplexity, Parallel advanced, Parallel basic, Firecrawl, Nimble, Tavily advanced, Tavily basic, You, Linkup standard, Parallel fast, Exa fast, Parallel turbo, Exa instant, Linkup fast, Brave. Every task is evaluated under two tool configurations: search only and search & fetch. When fetch is enabled, web_fetch goes through the same vendor under test as web_search— a search & fetch pair from one vendor, not search from one and fetch from another.

How to pick a provider.Start from whether fetch is the vendor's or yours. There is no single best search API. It is a tradeoff between grounded task completion, latency, tokens, and LLM cost.

  • Search & fetch. The agent searches and fetches through the same vendor. Read the search & fetch board.
  • Search only, bring your own fetch. The agent only calls web_search. Use this if fetch is your own scrape, browser, or another vendor. Read the search-only board.
  • Highest grounded completion. Read Task completion on the board that matches your tool setup.
  • Fastest loop. Read Median task time and Avg search time.
  • Leanest context. Read Median task tokens.
  • Lowest all-in cost. Read Median task cost. That is median LLM spend plus search/fetch list price per ticket. API list price is the unit rate of the search (and fetch) call.

Benchmark Data

Each chart ranks its own top eight; no single column decides the order. The tables list the full field alphabetically, with quality, latency, token usage and cost side by side.

Exa deep completes most tasks at 83.0%Perplexity is fastest at 21.5 sTinyFish costs least at $0.050

Top 8 per metric
Search & fetch coding-agent board, listed alphabetically
VendorEndpoint & configurationTask completionMedian task timeMedian task costAvg search timeMedian task tokensAPI list price
Exa autoPOST /search type=autoPOST /contents81.7 ± 1.123s$0.107$0.032 search+fetch$0.074 token cost1.19s27,433$0.007 / search$0.001 / fetch
Exa deepPOST /search type=deepPOST /contents83.0 ± 1.037s$0.127$0.057 search+fetch$0.070 token cost3.97s23,660$0.012 / search$0.001 / fetch
FirecrawlPOST /v2/searchPOST /v2/scrape76.0 ± 1.034s$0.082$0.028 search+fetch$0.055 token cost2.81s17,379$0.005 / search$0.0025 / fetch
Linkup standardPOST /v1/search depth=standardPOST /v1/fetch mode=standard48.3 ± 4.948s$0.147$0.038 search+fetch$0.108 token cost2.00s57,791$0.005 / search$0.005 / fetch
NimblePOST /v2/search search_depth=lite · full_content=false · focus=generalPOST /v2/extract formats=[markdown]60.3 ± 2.540s$0.054$0.054 token cost1.93s16,828$0.0011 / search
NimblePOST /v2/search search_depth=standard · full_content=false · focus=generalPOST /v2/extract formats=[markdown]45.0 ± 1.044s$0.078$0.078 token cost905ms28,586$0.005 / search
Parallel advancedPOST /v1/search mode=advancedPOST /v1/extract77.0 ± 1.033s$0.096$0.019 search+fetch$0.077 token cost3.11s27,092$0.005 / search$0.001 / fetch
Parallel basicPOST /v1/search mode=basicPOST /v1/extract76.0 ± 0.030s$0.107$0.023 search+fetch$0.084 token cost1.59s32,809$0.005 / search$0.001 / fetch
PerplexityPOST /search search_context_size=highPOST /search search_context_size=high77.7 ± 1.522s$0.083$0.025 search+fetch$0.058 token cost991ms20,062$0.005 / search$0.005 / fetch
Tavily advancedPOST /search search_depth=advancedPOST /extract extract_depth=advanced60.0 ± 2.046s$0.153$0.085 search+fetch$0.069 token cost3.41s26,269$0.016 / search$0.0032 / fetch
Tavily basicPOST /search search_depth=basicPOST /extract extract_depth=basic59.0 ± 1.732s$0.113$0.043 search+fetch$0.070 token cost1.50s27,405$0.008 / search$0.0016 / fetch
TinyFishGET api.search.tinyfish.aiGET api.fetch.tinyfish.ai format=markdown79.0 ± 2.024s$0.050$0.000 search+fetch$0.050 token cost1.32s12,844$0 / search$0 / fetch
YouPOST /v1/search extraction_mode=highlightsPOST /v1/contents55.0 ± 1.727s$0.119$0.027 search+fetch$0.092 token cost638ms42,806$0.005 / search$0.001 / fetch
YouPOST /v1/search extraction_mode=highlights · knowledge=corePOST /v1/contents54.0 ± 2.028s$0.117$0.026 search+fetch$0.091 token cost678ms37,453$0.005 / search$0.001 / fetch

Task completion is mean ± SD of 3 runs; n = 100 tasks. Median task time, median task cost, avg search time, and median task tokens are pooled across those same 3 runs.

Perplexity completes most tasks at 77.3%Perplexity is fastest at 18.3 sTinyFish costs least at $0.035

Top 8 per metric
Search-only coding-agent board, listed alphabetically
VendorEndpoint & configurationTask completionMedian task timeMedian task costAvg search timeMedian task tokensAPI list price
BravePOST /res/v1/llm/context38.0 ± 2.620s$0.082$0.025 search$0.057 token cost523ms14,103$0.005 / search
Exa fastPOST /search type=fast66.3 ± 1.520s$0.103$0.035 search$0.068 token cost626ms22,344$0.007 / search
Exa instantPOST /search type=instant61.3 ± 2.921s$0.105$0.035 search$0.070 token cost447ms22,423$0.007 / search
FirecrawlPOST /v2/search70.3 ± 1.528s$0.059$0.025 search$0.034 token cost2.87s7,456$0.005 / search
Linkup fastPOST /v1/search depth=fast43.3 ± 1.127s$0.100$0.025 search$0.075 token cost1.39s24,057$0.005 / search
NimblePOST /v2/search search_depth=lite · full_content=false · focus=general46.7 ± 3.833s$0.043$0.005 search$0.037 token cost3.15s8,713$0.0011 / search
NimblePOST /v2/search search_depth=standard · full_content=false · focus=general42.3 ± 2.525s$0.079$0.025 search$0.054 token cost878ms14,859$0.005 / search
Parallel fastPOST /v1/search mode=fast66.7 ± 1.522s$0.054$0.005 search$0.049 token cost953ms12,460$0.001 / search
Parallel turboPOST /v1/search mode=turbo64.7 ± 2.119s$0.059$0.005 search$0.054 token cost333ms14,130$0.001 / search
PerplexityPOST /search search_context_size=low77.3 ± 2.118s$0.059$0.025 search$0.034 token cost957ms8,765$0.005 / search
Tavily basicPOST /search search_depth=basic51.0 ± 5.326s$0.097$0.040 search$0.057 token cost1.72s16,299$0.008 / search
TinyFishGET api.search.tinyfish.ai59.3 ± 1.528s$0.035$0.000 search$0.035 token cost2.15s7,469$0 / search
YouPOST /v1/search extraction_mode=highlights41.3 ± 3.124s$0.097$0.025 search$0.072 token cost596ms20,333$0.005 / search
YouPOST /v1/search extraction_mode=highlights · knowledge=core38.3 ± 3.221s$0.096$0.025 search$0.071 token cost677ms20,786$0.005 / search

Task completion is mean ± SD of 3 runs; n = 100 tasks. Median task time, median task cost, avg search time, and median task tokens are pooled across those same 3 runs.

[02] methodology and metric definitions+

Real User Workflow in this Benchmark

We evaluate web search in a coding-agent workflow using realistic product tickets.

The agent receives a small existing httpx client and must patch it against the vendor's current official API documentation. The ticket describes the desired capability in natural language, but deliberately does not reveal the exact implementation details.

The ground-truth answer requires discovering an opaque detail from official documentation, such as:

  • an API path
  • a required header
  • a function or import name
  • another vendor-specific implementation identifier

Target documentation is enterprise SaaS documentation like Workday and SAP.

The benchmark is intentionally designed so that transforming ticket language into an API-looking string should not be enough. The agent should have to search.

The no-search condition is a necessity. If the model can solve a task by guessing the API or documentation structure, that task may still measure coding ability, but it no longer provides a clean measurement of web-search quality.

How the benchmark is built

Tasks are built from currently live official vendor documentation. Finding those pages is done by a script, not by the coding agent. The search path used here is a neutral third party that is not being evaluated in this benchmark.

  1. Discover. We start from a known list of official docs pages. A script then runs ordinary web searches for more pages on the same enterprise SaaS documentation sites. Anything not on those official sites is dropped. No LLM is in this step.
  2. Fetch. A second script downloads the full text of those pages into a local folder. Still no LLM. The Authoring Agent only sees these saved files later.
  3. Generate. An Authoring Agent reads only those stored pages and proposes product tickets: opaque gold token from the page, capability English that does not name it, and a starter that calls a real but wrong API. Each ticket has a minimum of three very similar configurations; disambiguation is necessary to find the right one.
  4. Verify. Automatic keep-gate rejects a ticket if it does not follow dataset guidelines such as naming the ground truth keywords in the ticket.
  5. Manual review. A human reads each accepted draft before it is promoted.

Before a task enters the dataset, it must pass an explicit no-search baseline. If the model can recover the ground-truth implementation from the ticket and starter alone, the task is excluded or rewritten.

The public 30-ticket set is on Hugging Face as openbenchmarks/OB-Code-Websearch. The boards on this page are scored on a held-out private set that is not distributed so that vendors and models cannot train and fit to the Benchmark. Use the public rows to inspect the format; scores on those rows are not comparable to the boards. Harness, scoring and vendor runners are in openbenchmarks-labs/web-search-for-coding-agents.

The model receives only the ticket and the starter code. It is run under one of two tool configurations:

[web_search]
[web_search, web_fetch]

Current limits:

Model: gpt-5.6-sol
Max turns: 32
Search budget: 5
Fetch budget: 5
Repeats: 3

Each vendor is run three times on the same locked ticket set, on both the search-only board and the search & fetch board, to account for model variability. The published task-completion score is the mean of those three runs, shown with the sample standard deviation. Median task time, median task cost, avg search time, and median task tokens are a single pooled number from the same three runs, not ± SD.

What each metric means

  • Task completion · mean share of tickets that pass, across three independent runs of the same 100-task set. A submission passes only if the submitted file compiles, all ground-truth tokens are present, each has an associated # source: URL, and that URL actually appeared in the search or fetch results from that run. A sufficiently informative search snippet is enough. The table reports mean ± SD of the three board pass rates.
  • Median task time · median end-to-end wall-clock time for the agent to finish the task, from the first turn until submission, pooled across all three repeats.
  • Avg search time · mean latency of each web_search tool call, pooled across all three repeats.
  • Median task tokens · median LLM prompt plus completion tokens per ticket, pooled across all three repeats.
  • API list price· PAYG dollar rate of one search call, and of one fetch call on the search & fetch board.
  • Median task cost · median LLM dollar cost per ticket (Braintrust estimated_cost() on the eval trace) plus median search/fetch API spend (call counts × list price), pooled across all three repeats.

Web search APIs for coding agents and developers

Can the model pass these coding tasks without web search?

No. The model-only baseline is 0. A pass requires a # source: URL that actually appeared in that run's search or fetch results. Without a search tool the model cannot provide that URL, and the ticket does not name the implementation detail.

What is the best web search API for AI coding agents in 2026?

Exa deep leads with fetch at 83.0% task completion on 100 tickets. Perplexity leads search-only at 77.3%. Compare task completion, latency, token usage, and cost for your workflow. Providers are listed alphabetically, with search-only and search-plus-fetch reported separately and no overall winner.

What is the best web search API for developers?

Exa deep leads with fetch at 83.0% task completion on 100 tickets. Perplexity leads search-only at 77.3%. Same board. Developers and coding agents use web search APIs to find implementation details in live official docs. This page compares quality, latency, token usage, and cost for that hard-retrieval job, with search+fetch and search-only reported separately.

Does this benchmark measure Context7, Ref, or MCP code search?

No. Those are documentation MCP / code-search tools. This board measures general web search APIs (Exa, Tavily, Brave, Parallel, Linkup, Firecrawl, TinyFish) used as search in a coding agent. They are a different product.

Why don't we use popular docs like Stripe?

The benchmark is designed against the model's reward hacking in the user workflow. Vendors with highly regular, developer-friendly documentation can be poor fits: in earlier Stripe tasks, the model often inferred the documentation slug or URL structure directly from the ticket and reached the correct implementation without performing meaningful retrieval. For example, when documentation followed a predictable pattern such as /changelog/dahlia/{kebab-case-ticket-title}, the model could construct the likely documentation location from the ticket itself. That tests the model's ability to exploit documentation naming conventions, not the quality of the search provider. Those tasks were removed from the active benchmark. Keep-gate: if a capable model can reliably solve the task without searching, the task does not belong in the primary search benchmark.

Web search tasks

This is the hard retrieval task on the web search benchmark.

Factual lookup: company news →

Multi-hop: multi-turn company search →

Best web search API comparisons

Best web search APIs for AI agents · Best web search APIs for developers · Best web search APIs for LLM apps & agents · Best web search APIs for LLM agents · Agentic search comparison & benchmark · Deep search comparison & benchmark · Web search APIs for AI & LLM developers · Web search APIs & MCPs for AI agents and developers · Search tools for AI agents · Search providers for LLM applications · AI search engines for agents · Free web search APIs for AI agents · Web search APIs for RAG · Independent web search API comparison · Best fast web search API · Best search and scrape API · Best scrape API for AI agents · Best search API for deep research agents · Best search API for company research · Best search API for sales agents · Best search API for coding documentation · Best web search API for grounding · Best search API for news

Best search APIs for coding agents: Best Web Search API for Coding Agents · Best Search Tool for Coding Agents · Best Search API for AI Coding Agents

Documentation lookup: Best Search API for Up-to-Date Docs · Best Technical Search API for AI Agents · Best Search and Fetch API for Coding Agents

By accuracy, speed, cost and tokens: Cheapest Search API for Coding Agents · Fastest Web Search API for Coding Agents · Most Accurate Web Search API for Coding Agents · Most Token-Efficient Web Search API for Coding Agents

Web search MCP: Best Web Search MCP for Claude Code · Best Web Search MCP for Cursor · Best Web Search MCP for Codex

All web search for coding agents comparisons →

[05] changelog+

Benchmarked Nimble lite and standard on 100 tasks across three repeats in both modes. Search uses full_content=false and focus=general; search + fetch calls POST /v2/extract separately with formats=[markdown].

Replaced Tavily Fast with Tavily Basic on the search-only board; evaluated 100 tasks across three repeats.

Re-evaluated TinyFish after updates were rolled out to their GA Fetch endpoint.