methodology
Methodology
How the cohort is built, how each route is called, what each metric means, and what these benchmarks cannot show.
The cohort
Cohort v1 is 50 US-English queries and 20 domains, declared before any call was made and frozen; a change means a new version, not an edit. The queries are the buyer vocabulary of the data-API market this site network covers (enrichment, SEO data, ad libraries, agent tooling), taken from a keyword research export on 14 September 2026. They deliberately mix head terms with monthly Google volume in the thousands and long-tail terms with no reported volume, because a keyword API’s coverage of the long tail is what an agent actually tests. The domains are the vendors compared across this site network plus treg.to. The cohort file is on the evidence page.
Routes
Every route is called through the Treg CLI so that credentials, retries and timing are handled the same way for every provider. A route is either a direct provider endpoint (DataForSEO, Serpstat, SE Ranking, SerpApi, Cloro, Moz, Google Ads Keyword Planner via the account’s own credentials) or Treg’s routed endpoint, which picks a provider per call and names the one that served. When a routed call is served by a provider that is also measured directly, the page says so; the two are the same data through two doors, not two datasets.
Routes we could not measure on the run date: Ahrefs (the account’s plan does not include API access), Majestic (no credential), Semrush (no credential). They are absent, not scored.
Settings held constant
Country US (Google location code 2840), language English, desktop, first page, depth 10, one call per query per provider, all providers run on the same day. Bulk keyword endpoints receive all 50 keywords in one request because that is how they are used; per-query endpoints receive one query per call. One worker per provider, so our own concurrency does not inflate any provider’s latency; providers do not run against each other on the same connection.
Metrics
Keyword coverage. For each provider: how many of the 50 keywords came back at all, how many with a numeric volume, how many as an explicit zero, and whether the provider distinguishes “no data” from “zero”. Agreement is reported against Google Ads Keyword Planner as the share of keywords within a factor of 2 (0.3 in log10) of Google’s average monthly figure, plus the median log ratio. Keyword Planner is the comparison point because it is what advertisers buy against, not because it is correct; it rounds to buckets and merges close variants.
SERP completeness. The reference for each query is a live rendered Google results page captured by Cloro at run time. For every other provider, the top-10 organic URLs (normalised: no scheme, no www, no trailing slash) are compared with the reference: overlap is the share of reference URLs present in the provider’s top 10, and “same first result” counts queries where position 1 matches. Results counts, calls that returned no data, and errors are reported separately. Serpstat’s SERP is a stored database, not a live fetch, so a query it has never indexed returns “keyword not found”; that is counted as no data, not as an error.
Backlink agreement. For each domain, each provider’s referring-domain count, backlink count and its own authority score. Spread is the ratio of the largest to the smallest referring-domain count across providers for that domain; rank correlation (Spearman) between providers is reported across the 20 domains. A larger count is not a more accurate count; no ground truth exists for backlinks and none is claimed.
Latency and cost. Latency is wall-clock from the CLI, including network, p50 and p95 per provider. Cost per call is the provider’s list price as read from the Treg catalog on the run date, except Treg routed calls, which report the amount actually charged. Rate limits hit during a run are recorded on the page and the affected call is retried after the run, with both attempts kept.
What these benchmarks cannot show
- Which provider has the most accurate volumes. There is no ground truth for search volume.
- Anything about your queries. Fifty queries in one market is a sample, and coverage on a different vocabulary or country will differ.
- Data freshness for stored-index providers, beyond what the run-to-run diff shows once there is more than one run.
- Provider quality at scale: rate limits, quotas and throughput at thousands of calls are not exercised by a 50-query cohort.
Dates
Each benchmark page shows the run date. Re-runs are published as separate dated runs; the page’s headline numbers always refer to the run named on the page, and old runs remain downloadable.
Exclusions
One competitor is excluded from this site under a private publishing policy. Comparisons are of selected routes and do not claim to cover the market.