benchmarks/company funding
04 · company funding

Company Funding Benchmark

What a funding data provider is asked for. Given a company domain, return that company's latest funding stage: the most recent round it raised, named correctly.

Two boards. Freshness measures rounds announced in the last 30 days: how fast a provider indexes news. Enrichment measures rounds older than that: how completely it has backfilled history. These reward opposite things, and a provider that leads one routinely trails the other, so they are ranked separately rather than averaged into a rank that describes neither. Each freshness cohort ages into enrichment once its window closes.

Vendors benchmarked. The 16 measured providers are Apollo, CompanyEnrich, Crunchbase, Crustdata, Exa, Explorium, Fiber, Firecrawl, Fundable, Harmonic, Ocean.io, Parallel, People Data Labs, PredictLeads, Seltz, and ZoomInfo. Several are measured on more than one endpoint, for 21 arms in total. Crunchbase and Harmonic are measured from reviewed exports mapped to the same company identities, so neither carries a latency or cost figure. PitchBook and Tracxn are private-market research platforms built for investor workflows and hence are not benched here.

Which board answers your question. Start from what consumes the data instead of the top of a ranking. Vendors are divided in 3 categories based on how they source data: Long Running Agent APIs research each company at request time, Web search APIs extract it from a search index, and GTM data providers return a stored record.

  • Freshness of Funding Data. Only accuracy matters in freshness on recent rounds. All three kinds of providers have comparable scores, so pick on price.
  • CRM Backfill Enrichment. Speed does not matter in a backfill enrichment job, so measure on accuracy and price. Read the enrichment board along with the funding fields returned, where the long running agents are the most accurate.
  • Enrichment at Scale. Price matters here the most. Compare correct stages per dollar across the web search APIs and GTM data providers. Long running agents are the most expensive and are not recommended for running enrichment at scale.
  • Live Agent Enrichment Workflow. Speed and accuracy matter the most here. Read median latency first, then correct when returned, and only among web search APIs and GTM providers. Long running agent APIs are significantly slower than the other two types of providers, and are not recommended for this usecase.

2026-08-26 added Firecrawl Spark 2 and rebenchmarked every vendor on the freshness board against a new snapshot, last run 2026-08-26. Full changelog →

open data + codeInputs, normalized provider outputs, and evaluation code are public in openbenchmarks-labs/company-funding.github →

Compare latest-stage accuracy, returned data, and speed.

Every column is a measured value. Column definitions, the matching policy, and what each number does not tell you are in [02] methodology.

Funding freshness: rounds announced in the last 30 days

Measured on 181 company observations across 3 dated snapshots: 2026-08 (2026-07-27 to 2026-08-26), 2026-08 (2026-07-16 to 2026-08-15) and 2026-07 (2026-06-16 to 2026-07-15). This is a rolling measurement: every vendor is re-run on a new snapshot of recent rounds, so these numbers change over time. Each snapshot above is scored on the rounds announced inside its own window.

benchmarks/funding/freshnessrolling
Funding providers compared on rounds announced in the last 30 days (window ending 2026-08-26)
ProviderEndpoint & configurationMeasured onLatest stage correctCorrect when returnedLatest stage returnedFunding fields returnedFields returnedMedian latencyEst. costOfficial docsCompany identified
Long Running Agent APIs
ExaPOST /agent/runseffort=medium · JSON schema100company measurements · 2 snapshots97.0%97.9%96.0%73.6%3.7 / 527,293 ms$10.00$0.10 / request (medium effort)Exa official docs 100.0%100/100 domains
FirecrawlPOST /v2/agentmodel=spark-2 · JSON schema · maxCredits capped per run50company measurements · 1 snapshot100.0%100.0%98.0%87.2%4.4 / 5102,203 ms$1.98$0.59 / 1,000 creditsFirecrawl official docs 100.0%50/50 domains
FirecrawlPOST /v2/agentmodel=spark-1-mini · JSON schema · maxCredits capped per run100company measurements · 2 snapshots96.0%96.9%97.0%94.4%4.7 / 592,998 ms$9.95$0.59 / 1,000 creditsFirecrawl official docs 100.0%100/100 domains
ParallelPOST /v1/tasks/runsprocessor=core · structured JSON output181company measurements · 3 snapshots95.0%96.0%95.6%86.7%4.3 / 562,142 ms$4.53$25 / 1,000 Task runsParallel official docs 100.0%181/181 domains
Web search APIs
ExaPOST /searchtype=instant · JSON schema100company measurements · 2 snapshots98.0%99.0%95.0%90.2%4.5 / 51,618 ms$0.70$0.007 / requestExa official docs 100.0%100/100 domains
ExaPOST /searchtype=deep-reasoning · JSON schema181company measurements · 3 snapshots95.6%96.0%96.7%84.8%4.2 / 59,916 ms$2.71$0.015 / requestExa official docs 100.0%181/181 domains
ParallelPOST /v1/responsesmodel=parallel · reasoning.effort=medium · text.format JSON schema100company measurements · 2 snapshots90.0%90.6%96.0%82.6%4.1 / 517,020 ms$5.00$50 / 1,000 Responses requests (medium reasoning)Parallel official docs 100.0%100/100 domains
SeltzPOST /v1/answerscope=news · response_format JSON100company measurements · 2 snapshots71.0%85.9%78.0%57.8%2.9 / 52,109 ms--Pricing not disclosedSeltz official docs 100.0%100/100 domains
SeltzPOST /v1/answerscope=companies · response_format JSON100company measurements · 2 snapshots22.0%32.8%67.0%62.6%3.1 / 52,118 ms--Pricing not disclosedSeltz official docs 100.0%100/100 domains
GTM data providers
ApolloGET /api/v1/organizations/enrich181company measurements · 3 snapshots59.7%69.5%85.1%70.1%3.5 / 5320 ms$3.44$0.0196 / modeled creditApollo official docs 97.2%176/181 domains
CompanyEnrichGET /companies/enrich181company measurements · 3 snapshots12.7%19.1%63.5%56.8%2.8 / 5334 ms$1.77$0.0098 / creditCompanyEnrich official docs 96.7%175/181 domains
Crunchbase (exported dataset)Crunchbase self-serve plan exportidentity-mapped export · not an API run81company measurements · 1 snapshot95.1%98.7%96.3%92.6%4.6 / 5----Not comparableCrunchbase (exported dataset) official docs 100.0%81/81 domains
CrustdataPOST /company/enrichexact_match=true · fields=[funding] · 25 domains/batch181company measurements · 3 snapshots81.2%90.1%89.5%71.3%3.6 / 51,537 ms$54.30$0.30 / response creditCrustdata official docs 96.1%174/181 domains
ExploriumPOST /v1/businesses/match + bulk_enrichfunding_and_acquisition enrichment181company measurements · 3 snapshots5.5%13.6%36.5%40.4%2.0 / 5715 ms$11.00$0.04 / request proxyExplorium official docs 44.8%81/181 domains
FiberPOST /v1/company-search181company measurements · 3 snapshots74.6%80.2%92.3%88.3%4.4 / 5615 ms$4.11$0.020 / creditFiber official docs 97.8%177/181 domains
FundableGET /api/v1/companydomain query parameter100company measurements · 2 snapshots84.0%92.3%91.0%91.0%4.5 / 51,300 ms$3.53$0.05 / creditFundable official docs 91.0%91/100 domains
Harmonic (exported dataset)Supplied Harmonic company exportidentity-audited mapping · not an API run81company measurements · 1 snapshot81.5%81.5%100.0%91.8%4.6 / 5----Pricing not disclosedHarmonic (exported dataset) official docs 100.0%81/81 domains
Ocean.ioPOST /v2/enrich/company181company measurements · 3 snapshots11.1%19.4%54.1%29.4%1.5 / 5575 ms$1.11$0.064 / creditOcean.io official docs 96.7%175/181 domains
People Data LabsGET /v5/company/enrich181company measurements · 3 snapshots13.3%19.5%65.2%51.4%2.6 / 5325 ms$17.38$0.10 / successful matchPeople Data Labs official docs 92.3%167/181 domains
PredictLeadsGET /api/v3/companies/{domain}/financing_events181company measurements · 3 snapshots68.0%82.4%81.8%64.3%3.2 / 5561 ms$4.83$0.04 / request · 100 freePredictLeads official docs 84.0%152/181 domains
ZoomInfogtm companies enrich --file10 domains/batch · 4 funding fields181company measurements · 3 snapshots72.4%85.0%84.5%85.8%4.3 / 5958 ms$18.10$0.10 / credit launch promo · $0.35 standardZoomInfo official docs 90.6%164/181 domains

Funding enrichment: rounds announced more than 30 days ago

The historical board. A vendor is scored on the rounds that were already more than 30 days old when it ran. Each row's own denominator is shown in the company identified column.

benchmarks/funding/company-funding-enrichment-v3live
Funding information providers compared on latest funding-stage data
ProviderEndpoint & configurationLatest stage correctCorrect when returnedLatest stage returnedFunding fields returnedFields returnedMedian latencyEst. costOfficial docsCompany identified
Long Running Agent APIs
ExaPOST /agent/runseffort=medium · JSON schema88.7%93.0%90.7%67.3%3.4 / 526,549 ms$30.00$0.10 / request (medium effort)Exa official docs 100.0%300/300 domains
FirecrawlPOST /v2/agentmodel=spark-1-mini · JSON schema · maxCredits capped per run92.3%92.3%100.0%93.3%4.7 / 5109,835 ms$29.86$0.59 / 1,000 creditsFirecrawl official docs 100.0%300/300 domains
FirecrawlPOST /v2/agentmodel=spark-2 · JSON schema · maxCredits capped per run87.3%89.9%95.3%84.2%4.2 / 5104,406 ms$11.87$0.59 / 1,000 creditsFirecrawl official docs 99.7%299/300 domains
ParallelPOST /v1/tasks/runsprocessor=core · structured JSON output90.0%92.3%95.4%84.9%4.3 / 553,684 ms$5.47$25 / 1,000 Task runsParallel official docs 100.0%219/219 domains
Web search APIs
ExaPOST /searchtype=deep-reasoning · JSON schema88.6%88.6%100.0%93.6%4.7 / 512,848 ms$3.29$0.015 / requestExa official docs 100.0%219/219 domains
ExaPOST /searchtype=instant · JSON schema87.0%89.0%97.0%89.0%4.5 / 51,507 ms$2.10$0.007 / requestExa official docs 100.0%300/300 domains
ParallelPOST /v1/responsesmodel=parallel · reasoning.effort=medium · text.format JSON schema89.0%90.1%94.3%81.6%4.1 / 515,930 ms$15.00$50 / 1,000 Responses requests (medium reasoning)Parallel official docs 100.0%300/300 domains
SeltzPOST /v1/answerscope=companies · response_format JSON54.7%59.9%91.3%80.7%4.0 / 51,969 ms--Pricing not disclosedSeltz official docs 100.0%300/300 domains
SeltzPOST /v1/answerscope=news · response_format JSON40.0%59.3%66.3%46.3%2.3 / 51,926 ms--Pricing not disclosedSeltz official docs 100.0%300/300 domains
GTM data providers
ApolloGET /api/v1/organizations/enrich49.8%63.7%78.1%64.3%3.2 / 5284 ms$4.16$0.0196 / modeled creditApollo official docs 96.8%212/219 domains
CompanyEnrichGET /companies/enrich46.1%55.2%83.6%69.3%3.5 / 5342 ms$2.15$0.0098 / creditCompanyEnrich official docs 97.7%214/219 domains
Crunchbase (exported dataset)Crunchbase self-serve plan exportidentity-mapped export · not an API run85.8%90.8%94.5%87.8%4.4 / 5----Not comparableCrunchbase (exported dataset) official docs 98.6%216/219 domains
CrustdataPOST /company/enrichexact_match=true · fields=[funding] · 25 domains/batch79.0%89.2%88.6%68.4%3.4 / 51,469 ms$65.70$0.30 / response creditCrustdata official docs 94.1%206/219 domains
ExploriumPOST /v1/businesses/match + bulk_enrichfunding_and_acquisition enrichment21.9%43.8%40.6%46.4%2.3 / 5697 ms$13.32$0.04 / request proxyExplorium official docs 53.4%117/219 domains
FiberPOST /v1/company-search84.9%88.2%96.3%90.0%4.5 / 5632 ms$4.98$0.020 / creditFiber official docs 97.3%213/219 domains
FundableGET /api/v1/companydomain query parameter60.3%86.1%69.7%68.7%3.4 / 51,370 ms$10.60$0.05 / creditFundable official docs 70.7%212/300 domains
Harmonic (exported dataset)Supplied Harmonic company exportidentity-audited mapping · not an API run71.2%72.9%97.7%90.0%4.5 / 5----Pricing not disclosedHarmonic (exported dataset) official docs 98.6%216/219 domains
Ocean.ioPOST /v2/enrich/company5.5%14.3%35.2%21.1%1.1 / 5551 ms$1.35$0.064 / creditOcean.io official docs 96.3%211/219 domains
People Data LabsGET /v5/company/enrich67.6%80.9%83.6%64.7%3.2 / 5276 ms$21.02$0.10 / successful matchPeople Data Labs official docs 95.9%210/219 domains
PredictLeadsGET /api/v3/companies/{domain}/financing_events37.9%70.1%53.4%39.0%1.9 / 5562 ms$5.84$0.04 / request · 100 freePredictLeads official docs 57.5%126/219 domains
ZoomInfogtm companies enrich --file10 domains/batch · 4 funding fields46.1%78.9%58.5%62.1%3.1 / 5972 ms$21.90$0.10 / credit launch promo · $0.35 standardZoomInfo official docs 76.7%168/219 domains
[02] methodology+

Two boards, split on how old the round is

Finding a round announced this week and holding a correct historical record are different capabilities, and a vendor can be strong at one and weak at the other. Averaging them into a single number hides that, so they are measured separately. The split is on the announcement date: rounds inside the trailing 30 days are freshness, everything older is enrichment.

Enrichment

Rounds announced more than 30 days ago. A durable historical set that grows each cycle as freshness snapshots age into it.

Freshness

Rounds announced in the trailing 30 days, sourced from current public funding news and verified before inclusion. Measured while the rounds are still recent, which is what exposes update lag.

Rolling windows

Each cycle builds a new dated freshness cohort and re-runs every vendor against it. Freshness results pool across the snapshots a vendor took part in, so its own case count is shown beside its score.

Ground Truth reviewed

Latest-stage Ground Truth is reviewed per company against a primary source. Blank Ground Truth is a pass-through case.

Source first, provider second

  1. Sample candidate companies from a time-windowed funding-transition index and independently research the announced event.
  2. Keep records with a direct company newsroom, company-issued wire announcement, or equivalent primary source; retain source URL, date, stage, amount, and evidence note.
  3. Build each freshness cohort from current public funding news inside that window, then require a verified company-domain identity before inclusion. A round with no verifiable announcement date is not placed in a window at all.
  4. Freeze the input list before any provider is scored. Candidate-source fields are never used as provider outputs.

Latest funding stage

  • Evaluate provider stages against the reviewed Ground Truth. The rubric handles documented equivalents, including Series suffixes within the same letter, PE/growth variants, strategic variants, and Seed Bridge/Pre-Series A.
  • Series letters remain strict: a Series B variant can match Series B, but Seed and Pre-Seed remain distinct. Vague labels such as “early stage” do not match a specific Ground Truth stage.
  • For non-Series stage disagreements, an exactly matching announced date or round amount can establish correctness.

Latest stage correct

  • correct latest stages ÷ the Ground Truth-reviewed companies that provider was measured on
  • Missing or incorrect returned stages lower this result. Blank Ground Truth passes every vendor answer; Undisclosed Ground Truth accepts an Undisclosed or blank vendor stage.
  • Each vendor is scored against the companies it was actually measured on, not a board-wide total. A vendor added later is measured on more of the enrichment cohort than one added earlier, and vendors join freshness at different snapshots, so denominators differ per row and are published alongside the score.
  • “Correct when returned” divides cases that both returned a stage and were correct by every case that returned a stage. The two conditions are separate: a blank answer can still be scored correct when Ground Truth is blank or Undisclosed, so counting raw correct answers here would exceed 100%. “Latest stage returned” measures whether a provider returned a stage at all.

Comparable inputs, provider-appropriate execution

  • API providers receive the same company domain through their documented company or funding endpoint. The natural-language arms receive one shared instruction naming the same company and domain, plus the same output schema. Crunchbase is measured from a self-serve-plan export matched to the frozen cohort, not from an API request.
  • Three vendors are measured on more than one endpoint, and each arm changes exactly one thing while the instruction and the output schema stay identical. Exa runs /search at deep-reasoning and at instant, plus its Agent API; Parallel runs the Task API and the Responses API; Seltz runs its companies and news scopes. Contract tests beside the runner assert that parity, so a difference in score is attributable to the varied parameter rather than to a different question.
  • ZoomInfo and Crustdata are called in batches rather than one company at a time, so their latency is the batch round trip recorded against each company in it, not a per-company request time. Read those two rows accordingly.
  • Raw responses, normalized fields, request latency, status, and failure reason are stored for every run. Only DNS/transient transport failures are retried; non-resolutions and rate-limit outcomes remain visible.
  • Harmonic is measured from a reviewed export mapped to the frozen identities, but has no comparable self-serve endpoint; it is excluded from latency and cost comparisons. PitchBook and Tracxn are private-market investor-research products, including PE workflows, rather than self-serve GTM sales enrichment APIs. Usage and pricing are shown only where a provider's billing basis is known.

Three ways to answer the same question

Both boards are subdivided by how a provider produces its answer, because a single ranked list quietly compares things that are not substitutes for each other. Every provider is asked the same question about the same companies and judged by the same policy, so the rows remain directly comparable. The grouping is a reading aid, not a separate scoring rule, and nothing is excluded or weighted differently.

  • Long Running Agent APIs dispatch an agent that researches each company at request time, reading sources and deciding what the latest round was. They are the slowest and most expensive per call, and they are priced on how much work they do.
  • Web search APIs query a search index and extract the answer from what comes back. Fast and cheap per call, with accuracy that tracks how well a round was covered on the open web.
  • GTM data providers return a stored record from a maintained company database. Latency and per-record cost are predictable, and coverage is set by the vendor's own ingestion rather than by the query.

The distinction is about mechanism, so Exa and Parallel each appear in two groups. Exa's Agent API and Parallel's Task API dispatch an agent that researches the company; Exa's two /search arms and Parallel's Responses API resolve the question against an index. The instruction and output schema sent to all of them are identical, so the endpoint is the only variable. The grouping also explains results a single list would make look contradictory, since an agent that researches on demand can lead on historical rounds and still be beaten on rounds announced last week.

What each column means

Every result is aligned to the same company domains. A latest stage counts as correct under the documented Ground Truth matching policy, which includes selected stage equivalents and exact date or amount evidence for non-Series disagreements. A missing provider stage lowers the correct-stage result, except that a blank vendor stage is accepted when Ground Truth is Undisclosed.

  • Measured on is that provider's own denominator: each provider/company/snapshot observation it was scored on, and how many snapshots those came from. A company repeated in a later snapshot counts again because it is a new query moment. Two rows with different values here were not measured on the same cohort.
  • Latest stage correct is the headline. Of the scoreable companies, the share where the provider named the right stage.
  • Correct when returned divides correct answers that actually returned a stage by every case where the provider returned one.
  • Funding fields returned counts how often latest stage, date, round amount, total raised, and round count were populated. It does not check whether the four non-stage fields are correct.
  • Official docs links each provider's own API, endpoint, or access documentation. These describe vendor configuration and do not influence the score.
  • Est. cost is modeled from observed billing units at published rates, scaled to the companies that board scored.
[03] what is not measured+
  • Whether the other funding fields are right. Only the latest stage is judged for correctness with announced date as fallback. Round amount, total raised, and round count are counted for presence only, so a provider can score well on funding fields returned while carrying wrong amounts.
  • Full funding history. One round per company is evaluated: the latest one. A provider with an excellent complete round history and a weak latest round scores badly here.
  • Investor, valuation, or board data. Not requested and not scored, even where a provider returns it.
  • Everything else the vendor sells. Most of these are broad platforms. Contact data, intent, technographics, and workflow tooling are different products. A result here says nothing about them.
  • Rows do not always contain the same measurement observations. Providers join at different times, so each is scored on its own denominator. Comparing two rows with very different Measured on counts compares performance on two different cohorts.
  • Cost is modeled, not invoiced. It is derived from observed billing units at published rates and scaled to each board. Plan minimums, negotiated rates, and credit rules will move real spend. Explorium's figure is an explicit request proxy because its per-endpoint credit burn is not public.
  • Latency is one runner on one network. Median request time is indicative, and batch endpoints are timed per batch rather than per company.
  • Reviewed exports are not timed runs. Crunchbase and Harmonic are mapped from supplied exports, so they carry no latency or cost and cannot be read as evidence of API behaviour.
  • Freshness is a rolling number. Each snapshot is a different set of companies. A change between snapshots can be a real shift in a provider, or a harder cohort.

Frequently asked questions

Where are the official API docs for the funding data providers in this benchmark?

The provider table links directly to official API, endpoint, product, or access documentation for every measured provider, including Crunchbase, Harmonic, Apollo, People Data Labs, PredictLeads, Exa, and Parallel. Those links document vendor access and capabilities; vendor claims do not determine benchmark scores.

Does Openbenchmarks verify funding events with official sources?

Yes. Ground Truth is independently reviewed against official company newsrooms, investor announcements, regulatory filings where applicable, and company-issued wire releases. Provider outputs are then scored against that frozen reference using identical company inputs.

What is the difference between the freshness board and the enrichment board?

They answer different questions and are ranked separately. Freshness measures rounds announced in the trailing 30 days, so it tells you how fast a provider indexes a new announcement. Enrichment measures rounds older than 30 days, so it tells you how completely a provider has backfilled history. A provider can lead one and trail the other, which is exactly why they are not averaged into a single rank.

Why are providers scored on different numbers of companies?

Because they were measured on different companies. Providers join the benchmark at different times, and each freshness snapshot is its own cohort that ages into enrichment once its window closes. Every score is therefore computed against that provider's own denominator, shown in the Measured on column. Using a board-wide denominator would silently penalise every provider that was present before a later one joined.

Which company funding data provider is the most accurate?

It depends on which board you need, so read the one that matches your workflow rather than a blended rank. The tables report latest stage correct as the headline, alongside correct when returned, which separates a provider that is often wrong from one that is often silent. Both are shown because a provider that returns nothing is not the same as a provider that returns a wrong stage, and the right trade-off depends on what consumes the data.

Are Crunchbase and Harmonic comparable to the API providers?

Only on correctness and coverage. Both are measured from reviewed exports mapped to the same company identities rather than from timed, billed endpoint calls, so neither carries a latency or cost figure. An export also has no query moment relative to an announcement, which is what a freshness score measures.

How is a latest funding stage decided to be correct?

Against a documented matching policy, judged by gpt-5.6 at medium reasoning effort with the same prompt for every provider. The policy accepts documented equivalents for same-letter Series variants, PE and growth, strategic investment, Seed Bridge and Pre-Series A, and crowdfunding and grant labels, and keeps Seed and Pre-Seed distinct. For non-Series disagreements an exact announced date or amount match can establish correctness. Every verdict carries a reason and a decision basis.

Can I reproduce these numbers?

Yes. The frozen inputs, the normalized provider outputs, the judge policy, and the evaluation code are public in the open data repository linked at the top of this page. Each freshness cohort publishes its own input list, so a snapshot is reproducible on its own rather than only as part of the combined cohort. Literal vendor HTTP response bodies are not redistributed.

Why are PitchBook and Tracxn not benchmarked here?

They are private-market research platforms built primarily for investor workflows, including private equity, rather than company funding enrichment APIs that a GTM system can call per domain. This benchmark measures the enrichment job, so a provider has to expose a way to ask about a specific company and get a structured answer back.

[05] changelog+
  • Added Firecrawl Spark 2 as a separate measured arm alongside Firecrawl Spark 1 Mini, using the same prompt, output schema, and agent endpoint so the model is the only changed variable.
  • Rebenchmarked every vendor on the freshness board against a new 50-company snapshot of recently announced funding rounds.
  • Six new measured arms across four vendors: the Parallel Responses API at medium reasoning effort, the Exa Search API with type=instant, the Exa Agent API, Seltz on its companies and news scopes, and Firecrawl's agent endpoint.
  • Three vendors are now measured on more than one endpoint, and each arm changes exactly one thing while the instruction and output schema stay identical: Exa across /search at deep-reasoning and instant plus its Agent API, Parallel across POST /v1/tasks/runs and POST /v1/responses, and Seltz across its two search scopes. Contract tests beside the runner assert that parity, so a score difference is attributable to the varied parameter rather than a different question.
  • The variable is priced very differently in each case. Exa charges $0.007 per instant search, $0.015 per deep-reasoning search and $0.10 per pinned agent run, which makes the accuracy-per-dollar curve inside one vendor directly readable. Parallel charges $50 per 1,000 Responses requests against $25 per 1,000 Task runs. Firecrawl bills dynamic credits per run, so its usage is measured from what each run reports rather than assumed.
  • Both boards are now subdivided by how each provider answers: long running agent APIs that dispatch an agent to research a company at request time, web search APIs that extract from an index, and GTM data providers that return a stored record. Every row is judged identically and stays directly comparable; the grouping stops a single ranked list from implying that a two-minute research agent and a 200ms database lookup are interchangeable purchases. Exa and Parallel each appear in two groups, because an agent endpoint and a search endpoint from the same vendor work differently.
  • The benchmark now publishes two boards instead of one blended ranking. Freshness measures rounds announced in the trailing 30 days, as dated snapshots. Enrichment measures rounds older than that. A provider that indexes fast and one that indexes deeply were previously averaged into a single number that described neither.
  • Each freshness snapshot ages into the enrichment cohort once its window closes, so enrichment grows every cycle and freshness stays a rolling measurement rather than a frozen one. Added the second freshness snapshot, scored by the same stage policy.
  • Vendors added after a freshness window has closed are measured on enrichment only for that window. Scoring them on a month-old snapshot would measure what they knew long after the announcement, which is not what a freshness score means.
  • Scores are now computed on per-provider denominators. Providers join at different times, so each is scored against the companies it was actually measured on, shown in the Measured on column. A board-wide denominator would have penalised every provider present before a late joiner arrived.
  • Estimated cost is scaled to the number of companies each board scored, rather than restating the original 300-domain figure on boards that measured fewer.
  • Corrected correct when returned, which could exceed 100%. The judge passes any vendor answer where Ground Truth is blank and accepts a blank answer where it is Undisclosed, so a case can be scored correct having returned no stage; the numerator now counts only cases that both returned and were correct. The headline correct-stage yield and the fill rate both divide by eligible cases and never changed.
  • Separately, fixed a counting error in funding field coverage. The metric only recognised the field names emitted by the original provider adapters (latest_date, round_count), so any provider scored through the later web-research schema, which emits latest_announced_on and funding_round_count, was capped at 3 of 5 fields no matter what it actually returned. Both spellings are now counted; coverage rose for the affected providers, and no correctness score, ranking, or judge verdict changed.
  • Added Harmonic as the 13th measured provider from a supplied 300-record company export, identity-audited against the frozen cohort.
  • Normalized its funding fields and judged every latest-stage answer with the same gpt-5.6 medium-reasoning policy as the rest of the benchmark: 222 / 300 correct (74.0%).
  • Harmonic is ranked for funding-stage correctness. Because this was an export rather than a timed, billed endpoint run, latency and cost remain intentionally blank.
  • Refreshed the 300-company funding Ground Truth through manual review.
  • Added the LLM-judged latest-stage snapshot, evaluated with gpt-5.6 at medium reasoning effort across all 13 providers.
  • Added an auditable judge decision and concise reason to every provider-stage result.
  • Added documented equivalence rules for same-letter Series variants, PE/growth, strategic-investment, Seed Bridge/Pre-Series A, and crowdfunding/grant labels.
  • Kept Seed and Pre-Seed distinct; for non-Series disagreements, an exact announced-date or amount match may establish correctness.
  • Blank Ground Truth passes every vendor answer; Undisclosed Ground Truth accepts an Undisclosed or blank vendor stage.
  • Rejudged the affected Series-plus and Seed/Pre-Seed cases, plus ZYT after its Ground Truth changed to Pre-IPO.