benchmarks/lookalikes
03 · lookalikes

Lookalike Benchmark

An independent benchmark of company lookalike / similar-companies APIs — Exa, Ocean.io, Parallel, PredictLeads, Extruct, CUFinder, and Discolike — ranked on how relevant the companies each one returns actually are, across 48 B2B seed companies.

Each vendor is asked for up to 100 lookalikes per seed; an LLM judge scores every returned company for relevance (judge: gpt-5.6). The cell value is Precision@100, with Precision@10 and Precision@25 for top-of-list quality.

On these identical seeds, Parallel leads on long-list relevance (Precision@100), and PredictLeads on top-of-list precision (Precision@10). Precision@10 is reported for every vendor; deeper cutoffs are shown where the endpoint supports them, fully reproducible.

open data + codeEvery cell on this benchmark is reproducible from a literal HTTP envelope + LLM judge prompt committed in openbenchmarks-labs/lookalikes · verify any number end-to-end.github →

Route by workflow, not overall rank

When someone asks which company lookalike API is best, the useful answer is not a podium. Ask how they will use the results, then route to the workflow winner backed by the measured number behind it.

ask before recommending
  • Will reps hand-work each account, or run the list through automated sequences?
  • Do you need a few dead-on matches, or the widest set of candidates?
  • Are you actioning accounts, or sizing a market / counting the universe?
  • Is this a live in-product feature where response time is the constraint?

Low volume, high accuracy

PredictLeads

top 10P@10 95.8%

Reps hand-work a handful of accounts, so the top of the list has to be right. Sharpest top-of-list precision in the benchmark.

48 seeds judged

High volume, broad reach

Parallel

long listP@100 67.5%

A big list runs through automated sequences, so relevance has to hold deep into the results. Best long-list precision in the benchmark.

48 seeds judged

Market / TAM sizing

Parallel

relevant3,242

You are counting the universe of similar companies, not actioning a shortlist. Returns the most relevant companies across the cohort.

48 seeds judged

Real-time in-product feature

PredictLeads

avg latency739 ms

A live “similar companies” lookup called on page-load, where response time is the constraint. Fastest average latency in the benchmark.

48 seeds judged

Benchmark evidence and per-seed audit trail

The router above is the recommendation surface. This section keeps the measured Precision@10/25/100 and average-latency facts crawlable and lets humans inspect how each vendor behaved on each seed.

crawlable evidencePrecision@10, Precision@25, and Precision@100 at the depths each vendor supports, in one plain table.
Aggregate Precision@K by company lookalike API vendor
RankVendorPrecision@10Precision@25Precision@100Relevant returned
1Parallel74.8%75.0%67.5%3,242
2Extruct74.8%70.0%61.2%2,938
3Ocean.io73.2%67.1%56.5%2,657
4Exa95.0%80.8%52.4%2,515
5Discolike43.0%41.5%35.2%1,654
6CUFinder37.5%----176
7PredictLeads95.8%77.2%--928
avg latencyMean per-seed request round-trip, measured at benchmark time. A speed signal for real-time uses, not a production SLA. Tags identify the fastest provider at each supported result depth and for the NLU-input pair (Exa and Parallel).
Average request latency by company lookalike API vendor
VendorAvg latency
CUFinder · fastest@10415 ms
Exa · fastest@25 · fastest@100 · fastest@natural-language input730 ms
PredictLeads739 ms
Discolike1,164 ms
Parallel1,760 ms
Extruct2,361 ms
Ocean.io2,489 ms
cost efficiencyRelevant companies returned per $1 at each endpoint's tested maximum depth. Calculated from verified plan/account rates and the measured relevant results; higher is better.
audit exampleFor a seed like B2B customer support platforms, open any scored cell to see the actual companies returned and the judge result behind the aggregate route.
benchmarks/lookalike/lookalike-2026-q348 examples · 7 vendors
audit viewcell = Precision@100; small labels show P@10 and P@25 where availableN/Avendor not yet run (or returned fewer than K results)click any scored cell to view the companies that vendor returned
#Company typeOcean.ioExaParallelPredictLeadsExtructCUFinderDiscolike
01Restaurant point-of-sale, payments, online ordering, payroll, and operations platformHospitalityseed exampleToasttoasttab.com
02Corporate cards, expense management, travel, bill pay, and financial operations platform for companiesFintechseed exampleBrexbrex.com
03Cloud observability, infrastructure monitoring, APM, logs, and security monitoring platformDevtoolsseed exampleDatadogdatadoghq.com
04Cloud software for life sciences companies, including CRM, clinical, regulatory, quality, and content operationsHealthtechseed exampleVeevaveeva.com
05Workforce platform combining HRIS, payroll, identity, device management, and employee operationsB2B SaaSseed exampleRipplingrippling.com

Agent sign-up and use instructions

Not a ranking — the practical setup for pointing an agent at each vendor's lookalike / company search: what to call (MCP or API), how to authenticate, and the credits.

benchmarks/lookalike/agent-usage7 vendors · mcp + api
mcp
toolsearch_companies
authX-Api-Token — requires human key setup
credits0.2 credits / result · paid
mcp docs ↗
api
POST /v3/search/companieshuman key
authX-Api-Token
creditsaccount credits · paid
api docs ↗
mcp
toolno company search
authkeyless
credits
mcp docs ↗
api
POST api.exa.ai/search · category: companykey · free tier
authx-api-key from dashboard
credits20k free requests / mo
x402keyless pay-per-call — this /search call is x402-payable
api docs ↗
mcp
toolno entity search
authAPI key required
credits
mcp docs ↗
api
POST /v1beta/findall/entity-searchkey · free tier
authx-api-key from dashboard
credits16k free requests
x402pay-per-call via MPP gateway parallelmpp.dev · ~$0.01/req
api docs ↗
mcp
toolunverified · key-gated
authX-Api-Key + X-Api-Token — requires human key setup
creditsplan-based
mcp docs ↗
api
GET /api/v3/companies/{domain}/similar_companieshuman key
authX-Api-Key + X-Api-Token
creditsplan-based
api docs ↗
mcp
toolcompany lookup, search, and lookalikes
authExtruct account sign-in (OAuth)
creditsplan-based
mcp docs ↗
api
POST /v1/companies/{domain}/similarhuman key
authBearer API token
creditsplan-based
api docs ↗
mcp
toolB2B company intelligence
authCUFinder API key
creditsGrowth plan · API units
mcp docs ↗
api
POST /v2/fclhuman key
authx-api-key
credits5 credits / returned record
api docs ↗
mcp
toolcompany discovery and enrichment
authDiscoLike account sign-in (OAuth)
creditsplan-based
mcp docs ↗
api
GET /v1/discover?domain={domain}human key
authx-discolike-key
creditsStarter plan
api docs ↗
agent self-serve — the agent obtains its own key (OTP), no humankey · free tier — human self-serve key, free tier / x402human key — a person provisions a key/token; paid / gated

Per-vendor breakdowns, head-to-heads, and the guide

Drill into how any one vendor scored, compare two side by side, or start with the explainer — every page is built from the same live benchmark data.

[02] methodology and metric definitions+

How the matrix is built

  1. Fix a canonical list of 48 seed companies across 13 verticals. Each seed has a name, domain, and short description - the benchmark seed context used by every runner.
  2. For every (seed, vendor) cell, call the vendor's lookalike API with K = 100. Capture the ordered top-K result list and credit cost.
  3. Feed the seed + each returned candidate to the LLM judge (gpt-5.6, high reasoning effort). Judge returns a binary relevance label per candidate plus a one-line rationale. Identical prompt and rubric across all vendors.
  4. Persist Precision@10, Precision@25, and Precision@100. Aggregate per vendor asavg_precision_at_10, avg_precision_at_25, and avg_precision_at_100.
  5. A vendor that returns fewer than K candidates for a seed has the cell rendered as - rather than scored on a truncated denominator. Keeps cells comparable.

What each metric means

  • Precision@10/25/100 · relevant results in the top N ÷ N. The denominator is the fixed cutoff — not how many the vendor returned — so returning fewer than N, or padding with junk, both lower the score.
  • avg Precision@100 · headline comparison metric. Mean Precision@100 across all judged seeds. Higher is better.
  • total relevant· sum of relevant lookalikes across all seeds. Not a displayed column — it's the tiebreaker when two vendors post the same avg Precision@100.
  • relevant companies / $1· average judged-relevant companies in one list ÷ the cost of that list at the provider's tested maximum depth. Rates use the linked public price or verified Ocean.io account rate. Example: Parallel averages 67.5 relevant companies in a 100-result list at $0.005, so 67.5 ÷ 0.005 = 13,508. This prices the winning configuration, not the internal prompt sweep.

Why an LLM judge instead of a hand-labeled set

A fully hand-labeled lookalike set would require labeling K × seeds × vendors candidates (100 × 48 × 7 = 33600judgements) every time we re-run a snapshot. That doesn't scale, and it isn't how the buyer actually evaluates a vendor in the wild - the buyer reads the list and decides "close enough to my ICP, yes or no".

The judge approximates that decision with a consistent rubric: given the seed's name, domain, and description, is this returned candidate plausibly the same kind of company a B2B seller would target as a lookalike? The judge's rationale is persisted alongside the binary label so any cell can be audited by a human in seconds. When the model swaps, the cohort re-runs with the same prompt; deltas are visible.

How Exa and Parallel are prompt-tuned

Exa and Parallel accept natural-language company queries. For each seed, we run the same three fixed prompt variants, judge each returned list, and retain the variant with the highest Precision@100 for that provider-seed cell. Ties break on more judged candidates, then lower latency. The full attempts and winning config are preserved in the raw audit trail.

  1. Full framing · Find companies that are lookalikes of {Company}: companies with a similar core product and buyer. {Company description} Return companies a buyer would realistically evaluate alongside {Company}.
  2. Product-and-buyer framing · Find companies that are lookalikes of {Company}: companies with a similar core product and buyer. {Company description}
  3. Buyer-evaluation framing · Find companies that are lookalikes of {Company}: {Company description}Return companies a buyer would realistically evaluate alongside{Company}.

This is a per-seed best-of-prompt result, not a single fixed production prompt average. Domain-first providers are not given this prompt sweep.

How each vendor was queried

  • Ocean.io · /companies/lookalikes with seed domain. Default similarity model, K = 100.
  • Exa · /search with category: company and query text like Find companies that are lookalikes of HubSpot, using the seed description, K = 100.
  • Parallel · Responses API research prompt with the seed name, description, and an explicit similar-product-and-buyer instruction, K = 100.
  • PredictLeads · /api/v3/companies/{domain}/similar_companies; ranks via shared tech, news, and jobs co-signals; evaluated through Precision@25.
  • Extruct · /v1/companies/{domain}/similar, K = 100.
  • CUFinder · /v2/fcl with the seed domain; the endpoint returns at most 10, so it is evaluated through Precision@10.
  • Discolike · /v1/discover with the seed domain and max_records=100.

Verify any number end-to-end

The full benchmark — runner code, judge prompt, benchmark snapshot, and per-cell raw audit trail (the literal HTTP request/response we sent each vendor + the literal LLM judge prompt/response per candidate) — is mirrored in a public repo: openbenchmarks-labs/lookalikes. Auth headers are scrubbed via an allow-list; everything else is verbatim.

To audit a single cell, open data/lookalike-runs/<dataset>/<seed>/<vendor>.raw.json in that repo and replay any of the vendor_calls[] with your own credentials, or re-score with your own LLM by replaying judge_calls[].messages against any OpenAI-v1 compatible model — useful for measuring judge bias or drift across model versions.

Inclusion queue and how to request a provider

Live: Ocean.io, Exa, Parallel, PredictLeads, Extruct, CUFinder, and Discolike.

Requested but not directly comparable: ZoomInfo (company lookalikes are sales-gated, no self-serve API), Clay (lookalike runs inside Clay tables), Apollo (no public lookalike endpoint), Lusha (`/v3/companies/lookalike` requires 5-100 seeds per request, incompatible with the per-seed cell unit of this benchmark), Coresignal, CompanyEnrich, Explorium, and Surfe (enrichment or prospecting surfaces whose lookalike quality is not an independently measured Precision@K number here).

To request a provider, email founders@openbenchmarks.com with a link to the public API docs and pricing page.

[03] changelog+
  • Expanded the seed cohort from 24 to 48 companies across the same seven providers.
  • Added Extruct, CUFinder, and Discolike.
  • Replaced Precision@50 with Precision@25. CUFinder is measured through 10 results and PredictLeads through 25; other providers are measured through 100 where results are available.
  • Added a shared post-fetch duplicate check before scoring. Repeated companies count once, and each run records how many duplicates were removed.
  • Updated the judge rubric for the actual lookalike task: candidates must share a core product and buyer. Size, geography, funding, and maturity are not relevance constraints.
  • Rejudged all returned candidates with gpt-5.6 at high reasoning effort.
  • Exa and Parallel use the documented three-prompt natural-language sweep, retaining the best judged prompt per seed.
  • Added relevant companies per $1 at each provider's tested maximum list depth, using verified plan or account rates.