benchmarks/multi-turn company search
multi-turn company search

Multi Turn Company Search Benchmark

What this benchmark measures. Each question combines three or four constraints - such as headquarters, investor backing, accelerator participation, founding period, or funding history. A fixed research agent plans multiple searches, follows evidence, and returns a structured company set. For each benchmark configuration, the agent's native web search is replaced with the vendor's search API. The result is scored against a hand-labelled canonical company set.

Vendors benchmarked. Brave Search, Exa Deep, Exa Instant, Firecrawl, Linkup Fast, Linkup Standard, Parallel Basic, Parallel Advanced, Seltz Companies, Google SERP through RapidAPI, and Tavily Advanced. Configurations are rows because two modes from one vendor can behave like different products.

Workflows this benchmark answers. Use it when choosing search infrastructure for automated account-list building, investor or accelerator portfolio discovery, multi-criteria market mapping, acquisition target research, or a research agent that must return a complete and defensible company set.

2026-08-22 published the first complete benchmark with a frozen, hand-labelled gold dataset and 2,970 agent runs across 11 search configurations and two research modes; last run 2026-08-22. Full changelog →

open runner + judgeProvider adapters, local run artifacts, and deterministic scoring are public in openbenchmarks-labs/multi-turn-company-search.github →public dataset10 public company-discovery questions and frozen gold labels on Hugging Face as openbenchmarks/OB-Company-Websearch. Board scores use a separate locked 45-question set.huggingface →

Multi-constraint company sets, scored over repeated agent runs.

Rows are ranked by mean F1, then mean precision. Quality metrics report mean ± sample standard deviation across 3 trials, with each trial aggregated over all 45 questions. Column definitions are in [02] methodology.

Web Search only

The agent can issue focused searches and read the returned titles, URLs, and snippets. It cannot fetch page text.

benchmarks/multi-turn/search-onlycomplete
Search API configurations compared without page fetching
ProviderEndpoint & configurationF1PrecisionRecallExact setMedian timeMean turnsCost / agent runMeasured on
ExaPOST /search type=deepexa-deep45.4 ± 2.083.2 ± 3.033.7 ± 1.41.5 ± 1.389.5 s7.4$0.65645 questions135 agent runs
ExaPOST /search type=instantexa-instant43.3 ± 1.082.6 ± 3.832.2 ± 0.53.0 ± 1.349.8 s7.3$0.60145 questions135 agent runs
ParallelPOST /v1/search mode=basicparallel-basic41.4 ± 3.084.7 ± 6.129.9 ± 2.62.2 ± 0.069.1 s7.3$1.00845 questions135 agent runs
LinkupPOST /v1/search depth=fastlinkup-fast41.1 ± 1.782.8 ± 1.030.3 ± 1.50.7 ± 1.364.2 s7.4$0.83745 questions135 agent runs
TavilyPOST /search search_depth=advancedtavily-advanced41.1 ± 2.383.7 ± 4.229.8 ± 1.72.2 ± 0.092.2 s7.4$0.91045 questions135 agent runs
ParallelPOST /v1/search mode=advancedparallel-advanced40.8 ± 1.186.7 ± 2.029.2 ± 1.32.2 ± 0.087.4 s7.3$0.61445 questions135 agent runs
LinkupPOST /v1/search depth=standardlinkup-standard40.6 ± 0.984.0 ± 7.929.7 ± 1.41.5 ± 2.672.3 s7.3$0.83745 questions135 agent runs
FirecrawlPOST /v2/searchfirecrawl30.4 ± 1.177.3 ± 4.420.7 ± 0.72.2 ± 0.075.0 s7.9$0.27845 questions135 agent runs
Brave SearchGET /res/v1/web/searchbrave28.0 ± 1.766.9 ± 8.519.3 ± 0.50.7 ± 1.343.5 s8.0$0.26945 questions135 agent runs
SeltzPOST /v1/search scope=companiesseltz-companies14.5 ± 0.940.0 ± 3.19.4 ± 0.70.0 ± 0.055.2 s7.3$1.50845 questions135 agent runs
SERP (RapidAPI)GET google-search74.p.rapidapi.comserp0.4 ± 0.60.7 ± 1.30.3 ± 0.40.0 ± 0.031.4 s7.8$0.10345 questions135 agent runs

F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor.

Web Search plus Web Fetch

The same agent can also fetch an exact URL returned by search. Provider-native extraction is used where available. Other providers use a custom HTTP/browser fetcher using playwright.

benchmarks/multi-turn/search-and-fetchcomplete
Search API configurations compared with page fetching enabled
ProviderEndpoint & configurationF1PrecisionRecallExact setMedian timeMean turnsCost / agent runMeasured on
ExaPOST /search type=deepexa-deep48.2 ± 2.189.4 ± 1.236.0 ± 2.22.2 ± 2.295.8 s7.5$0.63945 questions135 agent runs
ExaPOST /search type=instantexa-instant44.9 ± 0.985.9 ± 2.533.5 ± 0.95.2 ± 1.352.5 s7.5$0.60645 questions135 agent runs
ParallelPOST /v1/search mode=advancedparallel-advanced42.4 ± 1.189.2 ± 2.130.0 ± 1.32.2 ± 0.083.9 s7.4$0.59945 questions135 agent runs
LinkupPOST /v1/search depth=standardlinkup-standard42.0 ± 1.890.7 ± 2.030.5 ± 2.03.0 ± 3.481.0 s7.5$0.84645 questions135 agent runs
ParallelPOST /v1/search mode=basicparallel-basic42.0 ± 4.087.0 ± 7.630.2 ± 3.23.0 ± 1.371.7 s7.2$0.94945 questions135 agent runs
TavilyPOST /search search_depth=advancedtavily-advanced41.0 ± 1.389.4 ± 6.029.1 ± 0.72.2 ± 0.092.6 s7.3$0.81345 questions135 agent runs
LinkupPOST /v1/search depth=fastlinkup-fast39.9 ± 1.385.3 ± 3.628.6 ± 1.30.7 ± 1.368.0 s7.3$0.82245 questions135 agent runs
FirecrawlPOST /v2/searchfirecrawl33.2 ± 2.183.3 ± 0.822.7 ± 1.81.5 ± 1.382.1 s7.9$0.29445 questions135 agent runs
Brave SearchGET /res/v1/web/searchbrave29.4 ± 1.673.5 ± 5.820.4 ± 0.91.5 ± 1.345.1 s8.0$0.27345 questions135 agent runs
SeltzPOST /v1/search scope=companiesseltz-companies16.3 ± 1.549.5 ± 2.410.2 ± 1.20.0 ± 0.060.1 s7.5$1.53045 questions135 agent runs
SERP (RapidAPI)GET google-search74.p.rapidapi.comserp0.0 ± 0.00.0 ± 0.00.0 ± 0.00.0 ± 0.033.4 s7.8$0.10345 questions135 agent runs

F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor.

[02] methodology +

45 company-discovery questions with hand labelled answer sets

Questions were selected from broad intersections in a frozen company census. Twenty-four questions have three constraints and twenty-one have four. The published gold release contains 375 canonical question-company memberships, with between two and thirty-seven valid companies per question.

How the dataset was created

The dataset was hand labelled. Reviewers checked companies against every constraint in each question, resolved names and domains to canonical company identities, and manually assembled the complete expected answer set. The resulting question-to-company memberships were then frozen as the gold set before benchmark scoring.

A separate 10-question public search-only set, including its frozen reference companies, is available on Hugging Face as openbenchmarks/OB-Company-Websearch. It can be used to inspect the schema and run the open harness, but scores computed on it are not comparable to this board, which uses the separate locked 45-question set. The runner and offline judge are in openbenchmarks-labs/multi-turn-company-search.

45 questions

Investor, accelerator, geography, founding-era, and funding constraints.

3 agent runs per question

Independent stochastic agent runs for every provider and question.

2 modes

Search-only and search-plus-fetch are evaluated separately.

2,970 agent runs

Every completed agent run is weighted equally in the reported averages.

One model and one prompt across providers

  • gpt-5.6-sol, medium reasoning effort.
  • Maximum 8 model turns and 14 searches.
  • Maximum two searches per turn and ten results per search.
  • The final response follows one strict JSON schema with company name, domain, cited URLs, and evidence.
  • The agent is instructed not to use prior knowledge as evidence and must complete at least one successful provider search.

The search provider is the comparison variable

  • The question, system prompt, model, reasoning effort, budgets, result limit, output schema, and agent-run count are held constant.
  • The agent chooses its own focused queries, so trajectories and search-call counts can differ after provider results diverge.
  • Provider responses are normalized into a shared title, URL, and snippet shape before the agent sees them.
  • Search-only is the clean provider comparison. Search-plus-fetch measures the provider's search and content-retrieval stack where a native fetch endpoint exists.

Deterministic set comparison

Precision

true positives ÷ all returned companies, reported as the three-trial mean ± SD.

Recall

true positives ÷ all gold companies, reported as the three-trial mean ± SD.

F1

The harmonic mean of precision and recall per agent run, reported as the three-trial mean ± SD.

Exact set

The share of agent runs with no false positives and no false negatives, reported as the three-trial mean ± SD.

Returned names and domains are resolved to canonical company identities before scoring. Human resolution overrides are versioned; unresolved names remain distinct raw predictions and therefore cannot receive a true-positive match by string coincidence alone.

Time, turns, and cost travel with quality

  • Median time is the median end-to-end wall-clock time across all agent runs for each vendor.
  • Mean model turns records how many model calls the provider's results caused.
  • Cost per agent run combines measured or list-priced vendor usage with model-token cost.
  • Every raw model response, vendor request, vendor response, normalized result, fetch, parse, and resolution is retained for replay.

Three agent runs per provider and question

Search agents are stochastic: the same question can produce different queries and company sets. Every provider receives three independent agent runs. For each quality metric, each trial is first averaged over all 45 questions. The table then reports the mean and sample standard deviation of those three trial-level values rather than selecting the best attempt. Standard deviation is shown in percentage points (pp), describes trial-to-trial variability, and is not a confidence interval.

What is not measured here

This not a universal web-search score

  • It does not measure one-shot fact lookup, breaking-news freshness, coding, academic literature, shopping, local search, or consumer navigation.
  • It does not measure Claude or another answer model.
  • Completed agent runs measure successful benchmark execution. They do not estimate provider uptime or rate-limit reliability under production traffic.
  • Search-plus-fetch does not isolate search alone because native content extraction is part of some provider configurations. Use the search-only board for that comparison.

Frequently asked questions

What is multi-turn web search?

Multi-turn web search requires an agent to plan and run more than one focused search, combine evidence from different results, and return an answer that satisfies several constraints at once. In this benchmark, every question asks for the complete set of companies matching three or four conditions rather than one isolated fact.

What does the Multi Turn Company Search Benchmark measure?

It measures whether a fixed research agent can use a search provider to recover the correct set of companies for each multi-constraint question. The benchmark reports precision, recall, F1, exact-set accuracy, cited-return rate, end-to-end latency, model turns, and cost.

Which web search APIs are included?

The first run includes Brave Search, Exa Deep, Exa Instant, Firecrawl, Linkup Fast, Linkup Standard, Parallel Basic, Parallel Advanced, Seltz Companies, a Google SERP API, and Tavily Advanced. Each endpoint and configuration is reported as a separate measured row.

How are providers scored?

Every returned company is resolved to a canonical company identity and compared with a human-reviewed gold set. A correct member is a true positive, an extra company is a false positive, and a missed gold company is a false negative. Precision, recall, and F1 are calculated per agent run and averaged across all questions and three independent agent runs.

Why does the benchmark report both precision and recall?

A company research answer can fail in two opposite ways. It can pad the result with companies that do not satisfy every constraint, lowering precision, or omit valid companies, lowering recall. F1 summarizes both, while exact-set accuracy records the stricter case where the returned set matches the reviewed set exactly.

What is the difference between search-only and search-plus-fetch?

Search-only exposes provider search results and snippets to the agent but no page-reading tool. Search-plus-fetch also lets the agent open an exact URL returned by search. The two modes are reported separately because page retrieval can change both answer quality and cost.

Does this benchmark test Claude web search?

No. The first run holds GPT-5.6 Sol at medium reasoning effort constant as the research agent. The benchmark compares search-provider configurations under that fixed agent. A Claude-powered run would require rerunning every provider because the model chooses its own searches and follow-up trajectory.

Is this the same as a factual question-answering search benchmark?

No. A factual question-answering benchmark usually asks for one short answer and scores final-answer accuracy. This benchmark asks for a complete entity set under multiple constraints, so it can directly measure false positives, false negatives, precision, recall, F1, and exact-set match.

Can these results identify the best search API for every workflow?

No. The dataset is specifically about multi-constraint company discovery. It does not measure general consumer search, news freshness, coding research, academic research, citation entailment, or performance with a different answer model.

[05] changelog+
  • Published the frozen, hand-labelled 45-question dataset and its human-reviewed canonical company sets for deterministic precision, recall, F1, and exact-set scoring.
  • Locked the agent contract, strict output schema, provider configurations, and three-agent-run evaluation design before scoring.
  • Separated search-only from search-plus-fetch so raw search quality can be read independently from page retrieval.
  • Completed coverage for all 11 configurations: 45 questions, three agent runs per question and mode, and 2,970 successful agent runs.