45 questions
Investor, accelerator, geography, founding-era, and funding constraints.
What this benchmark measures. Each question combines three or four constraints - such as headquarters, investor backing, accelerator participation, founding period, or funding history. A fixed research agent plans multiple searches, follows evidence, and returns a structured company set. For each benchmark configuration, the agent's native web search is replaced with the vendor's search API. The result is scored against a hand-labelled canonical company set.
Vendors benchmarked. Brave Search, Exa Deep, Exa Instant, Firecrawl, Linkup Fast, Linkup Standard, Parallel Basic, Parallel Advanced, Seltz Companies, Google SERP through RapidAPI, and Tavily Advanced. Configurations are rows because two modes from one vendor can behave like different products.
Workflows this benchmark answers. Use it when choosing search infrastructure for automated account-list building, investor or accelerator portfolio discovery, multi-criteria market mapping, acquisition target research, or a research agent that must return a complete and defensible company set.
2026-08-22 published the first complete benchmark with a frozen, hand-labelled gold dataset and 2,970 agent runs across 11 search configurations and two research modes; last run 2026-08-22. Full changelog →
Rows are ranked by mean F1, then mean precision. Quality metrics report mean ± sample standard deviation across 3 trials, with each trial aggregated over all 45 questions. Column definitions are in [02] methodology.
The agent can issue focused searches and read the returned titles, URLs, and snippets. It cannot fetch page text.
| Provider | Endpoint & configuration | F1 | Precision | Recall | Exact set | Median time | Mean turns | Cost / agent run | Measured on |
|---|---|---|---|---|---|---|---|---|---|
| Exa | POST /search type=deepexa-deep | 45.4 ± 2.0 | 83.2 ± 3.0 | 33.7 ± 1.4 | 1.5 ± 1.3 | 89.5 s | 7.4 | $0.656 | 45 questions135 agent runs |
| Exa | POST /search type=instantexa-instant | 43.3 ± 1.0 | 82.6 ± 3.8 | 32.2 ± 0.5 | 3.0 ± 1.3 | 49.8 s | 7.3 | $0.601 | 45 questions135 agent runs |
| Parallel | POST /v1/search mode=basicparallel-basic | 41.4 ± 3.0 | 84.7 ± 6.1 | 29.9 ± 2.6 | 2.2 ± 0.0 | 69.1 s | 7.3 | $1.008 | 45 questions135 agent runs |
| Linkup | POST /v1/search depth=fastlinkup-fast | 41.1 ± 1.7 | 82.8 ± 1.0 | 30.3 ± 1.5 | 0.7 ± 1.3 | 64.2 s | 7.4 | $0.837 | 45 questions135 agent runs |
| Tavily | POST /search search_depth=advancedtavily-advanced | 41.1 ± 2.3 | 83.7 ± 4.2 | 29.8 ± 1.7 | 2.2 ± 0.0 | 92.2 s | 7.4 | $0.910 | 45 questions135 agent runs |
| Parallel | POST /v1/search mode=advancedparallel-advanced | 40.8 ± 1.1 | 86.7 ± 2.0 | 29.2 ± 1.3 | 2.2 ± 0.0 | 87.4 s | 7.3 | $0.614 | 45 questions135 agent runs |
| Linkup | POST /v1/search depth=standardlinkup-standard | 40.6 ± 0.9 | 84.0 ± 7.9 | 29.7 ± 1.4 | 1.5 ± 2.6 | 72.3 s | 7.3 | $0.837 | 45 questions135 agent runs |
| Firecrawl | POST /v2/searchfirecrawl | 30.4 ± 1.1 | 77.3 ± 4.4 | 20.7 ± 0.7 | 2.2 ± 0.0 | 75.0 s | 7.9 | $0.278 | 45 questions135 agent runs |
| Brave Search | GET /res/v1/web/searchbrave | 28.0 ± 1.7 | 66.9 ± 8.5 | 19.3 ± 0.5 | 0.7 ± 1.3 | 43.5 s | 8.0 | $0.269 | 45 questions135 agent runs |
| Seltz | POST /v1/search scope=companiesseltz-companies | 14.5 ± 0.9 | 40.0 ± 3.1 | 9.4 ± 0.7 | 0.0 ± 0.0 | 55.2 s | 7.3 | $1.508 | 45 questions135 agent runs |
| SERP (RapidAPI) | GET google-search74.p.rapidapi.comserp | 0.4 ± 0.6 | 0.7 ± 1.3 | 0.3 ± 0.4 | 0.0 ± 0.0 | 31.4 s | 7.8 | $0.103 | 45 questions135 agent runs |
F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor.
The same agent can also fetch an exact URL returned by search. Provider-native extraction is used where available. Other providers use a custom HTTP/browser fetcher using playwright.
| Provider | Endpoint & configuration | F1 | Precision | Recall | Exact set | Median time | Mean turns | Cost / agent run | Measured on |
|---|---|---|---|---|---|---|---|---|---|
| Exa | POST /search type=deepexa-deep | 48.2 ± 2.1 | 89.4 ± 1.2 | 36.0 ± 2.2 | 2.2 ± 2.2 | 95.8 s | 7.5 | $0.639 | 45 questions135 agent runs |
| Exa | POST /search type=instantexa-instant | 44.9 ± 0.9 | 85.9 ± 2.5 | 33.5 ± 0.9 | 5.2 ± 1.3 | 52.5 s | 7.5 | $0.606 | 45 questions135 agent runs |
| Parallel | POST /v1/search mode=advancedparallel-advanced | 42.4 ± 1.1 | 89.2 ± 2.1 | 30.0 ± 1.3 | 2.2 ± 0.0 | 83.9 s | 7.4 | $0.599 | 45 questions135 agent runs |
| Linkup | POST /v1/search depth=standardlinkup-standard | 42.0 ± 1.8 | 90.7 ± 2.0 | 30.5 ± 2.0 | 3.0 ± 3.4 | 81.0 s | 7.5 | $0.846 | 45 questions135 agent runs |
| Parallel | POST /v1/search mode=basicparallel-basic | 42.0 ± 4.0 | 87.0 ± 7.6 | 30.2 ± 3.2 | 3.0 ± 1.3 | 71.7 s | 7.2 | $0.949 | 45 questions135 agent runs |
| Tavily | POST /search search_depth=advancedtavily-advanced | 41.0 ± 1.3 | 89.4 ± 6.0 | 29.1 ± 0.7 | 2.2 ± 0.0 | 92.6 s | 7.3 | $0.813 | 45 questions135 agent runs |
| Linkup | POST /v1/search depth=fastlinkup-fast | 39.9 ± 1.3 | 85.3 ± 3.6 | 28.6 ± 1.3 | 0.7 ± 1.3 | 68.0 s | 7.3 | $0.822 | 45 questions135 agent runs |
| Firecrawl | POST /v2/searchfirecrawl | 33.2 ± 2.1 | 83.3 ± 0.8 | 22.7 ± 1.8 | 1.5 ± 1.3 | 82.1 s | 7.9 | $0.294 | 45 questions135 agent runs |
| Brave Search | GET /res/v1/web/searchbrave | 29.4 ± 1.6 | 73.5 ± 5.8 | 20.4 ± 0.9 | 1.5 ± 1.3 | 45.1 s | 8.0 | $0.273 | 45 questions135 agent runs |
| Seltz | POST /v1/search scope=companiesseltz-companies | 16.3 ± 1.5 | 49.5 ± 2.4 | 10.2 ± 1.2 | 0.0 ± 0.0 | 60.1 s | 7.5 | $1.530 | 45 questions135 agent runs |
| SERP (RapidAPI) | GET google-search74.p.rapidapi.comserp | 0.0 ± 0.0 | 0.0 ± 0.0 | 0.0 ± 0.0 | 0.0 ± 0.0 | 33.4 s | 7.8 | $0.103 | 45 questions135 agent runs |
F1, precision, recall, and exact-set accuracy are percentages reported as mean ± sample SD across three independent runs; each run aggregates all 45 questions. SD is measured in percentage points. Median time is the median end-to-end time across all runs for each vendor.
Questions were selected from broad intersections in a frozen company census. Twenty-four questions have three constraints and twenty-one have four. The published gold release contains 375 canonical question-company memberships, with between two and thirty-seven valid companies per question.
The dataset was hand labelled. Reviewers checked companies against every constraint in each question, resolved names and domains to canonical company identities, and manually assembled the complete expected answer set. The resulting question-to-company memberships were then frozen as the gold set before benchmark scoring.
A separate 10-question public search-only set, including its frozen reference companies, is available on Hugging Face as openbenchmarks/OB-Company-Websearch. It can be used to inspect the schema and run the open harness, but scores computed on it are not comparable to this board, which uses the separate locked 45-question set. The runner and offline judge are in openbenchmarks-labs/multi-turn-company-search.
Investor, accelerator, geography, founding-era, and funding constraints.
Independent stochastic agent runs for every provider and question.
Search-only and search-plus-fetch are evaluated separately.
Every completed agent run is weighted equally in the reported averages.
true positives ÷ all returned companies, reported as the three-trial mean ± SD.
true positives ÷ all gold companies, reported as the three-trial mean ± SD.
The harmonic mean of precision and recall per agent run, reported as the three-trial mean ± SD.
The share of agent runs with no false positives and no false negatives, reported as the three-trial mean ± SD.
Returned names and domains are resolved to canonical company identities before scoring. Human resolution overrides are versioned; unresolved names remain distinct raw predictions and therefore cannot receive a true-positive match by string coincidence alone.
Search agents are stochastic: the same question can produce different queries and company sets. Every provider receives three independent agent runs. For each quality metric, each trial is first averaged over all 45 questions. The table then reports the mean and sample standard deviation of those three trial-level values rather than selecting the best attempt. Standard deviation is shown in percentage points (pp), describes trial-to-trial variability, and is not a confidence interval.
Multi-turn web search requires an agent to plan and run more than one focused search, combine evidence from different results, and return an answer that satisfies several constraints at once. In this benchmark, every question asks for the complete set of companies matching three or four conditions rather than one isolated fact.
It measures whether a fixed research agent can use a search provider to recover the correct set of companies for each multi-constraint question. The benchmark reports precision, recall, F1, exact-set accuracy, cited-return rate, end-to-end latency, model turns, and cost.
The first run includes Brave Search, Exa Deep, Exa Instant, Firecrawl, Linkup Fast, Linkup Standard, Parallel Basic, Parallel Advanced, Seltz Companies, a Google SERP API, and Tavily Advanced. Each endpoint and configuration is reported as a separate measured row.
Every returned company is resolved to a canonical company identity and compared with a human-reviewed gold set. A correct member is a true positive, an extra company is a false positive, and a missed gold company is a false negative. Precision, recall, and F1 are calculated per agent run and averaged across all questions and three independent agent runs.
A company research answer can fail in two opposite ways. It can pad the result with companies that do not satisfy every constraint, lowering precision, or omit valid companies, lowering recall. F1 summarizes both, while exact-set accuracy records the stricter case where the returned set matches the reviewed set exactly.
Search-only exposes provider search results and snippets to the agent but no page-reading tool. Search-plus-fetch also lets the agent open an exact URL returned by search. The two modes are reported separately because page retrieval can change both answer quality and cost.
No. The first run holds GPT-5.6 Sol at medium reasoning effort constant as the research agent. The benchmark compares search-provider configurations under that fixed agent. A Claude-powered run would require rerunning every provider because the model chooses its own searches and follow-up trajectory.
No. A factual question-answering benchmark usually asks for one short answer and scores final-answer accuracy. This benchmark asks for a complete entity set under multiple constraints, so it can directly measure false positives, false negatives, precision, recall, F1, and exact-set match.
No. The dataset is specifically about multi-constraint company discovery. It does not measure general consumer search, news freshness, coding research, academic research, citation entailment, or performance with a different answer model.