Scraping cookbook · method

How the numbers are counted

Every figure on a Cookbook page is produced the same way, from the same counters that bill customers. This page is the definition.

What counts as a success

Crawlbase records one row per request. A request is a success when Crawlbase answered with cb_status 200, 207, 227 or 228 and the site itself answered with 200, 201, 204, a redirect (301, 302, 303, 307 or 308), 404 or 410. Everything else is a failure.

A 404 on a product the site has removed therefore counts as a success: the request reached the site and the site answered. A block page, a challenge that could not be cleared, a timeout or an internal error counts as a failure. The success rate is successes divided by successes plus failures.

Which requests are counted

All of them. Every request from every account, across the Crawling API, the Crawler and the Smart AI Proxy, in one calendar month. No sampling, no exclusion of accounts or of retries. The month is printed on every page; the numbers are refreshed when the next month closes.

Sites are keyed by their registrable name without the ending, the way the counters store them, so idealo.de, idealo.es and idealo.co.uk are one site. The domain shown on a page is the one accounts fetch most.

What browser share means

Every request is made with either a normal token or a JavaScript token. Browser share is the share of successful requests made with the JavaScript token. It says how the accounts chose to fetch the site, not what the site strictly requires; the recipe says that. "No" is under 5%, "Yes" is over 95%, anything between is shown as the share.

Tiers and credits

Every site sits in a pricing tier. A standard page costs one credit; other tiers multiply that: moderate 1.5, complex 2, complex I 3, complex II 4.5, extreme 9, extreme I 13.5, extreme II 20. The tier shown is the one in force on the day the page was generated. A JavaScript request costs the tier price plus the browser surcharge described on the pricing page.

How a recipe is written

A recipe comes from what the accounts fetching the site actually send and get back: the URL patterns, the country and token that succeed, the statuses the site answers with, the latency. Four recipes were also written by hand. Nothing on a page is a guess; every sentence names the number behind it. No customer is ever named or identifiable.

Refresh

The numbers are recomputed from the full calendar month after it closes and every page carries the month it shows. A site that falls below the bar loses its page until it clears it again.

Which sites get a page

A site appears in the Cookbook when three or more accounts fetched it with at least a thousand successes in the month. Hosting and CDN infrastructure, banks and payment brands, schools, and people-search and background-check services never get a page, whatever the numbers say.

Site names identify the pages a recipe fetches. Crawlbase is not affiliated with any site listed. Use recipes within the acceptable use policy.Measured numbers from August 2026, refreshed monthly