Tony Wang17 min readHow Much of the Web Runs Anti-Bot? We Scanned the Top 1,000,000 Sites
We scanned the top 1,000,000 sites: 53.5% of the reachable web runs a managed anti-bot or WAF — and, surprisingly, the busiest sites run the least.
Most "I'll just scrape it" projects do not die on parsing the HTML. They die when the page turns out to sit behind Cloudflare or DataDome and needs a completely different approach than the one you built. So we measured how common that actually is — twice. We scanned the whole Tranco top 1,000,000 to get the web-scale picture, and then hand-picked 1,005 real, public sites across 28 categories — the kinds of sites people actually scrape — and ran the full transport fleet against them. The short version: just over half of the reachable web runs managed bot defense, it is concentrated in one vendor, the busiest sites run the least of it, and the difficulty is wildly uneven.
The full, searchable dataset — every site, filterable by vendor, difficulty and category — lives in the Anti-Bot Adoption Index.
How much of the web runs an anti-bot or WAF?
Across the full top 1,000,000, we reached 818,614 sites (the rest were unresolvable, timed out, or blocked our datacenter vantage outright — about 18%, excluded from the percentages). Of the ones we reached, 53.5% expose a managed WAF, bot-management, or access-control vendor — not counting a CAPTCHA widget or rate-limiting alone:
| Protection type | Share of reachable |
|---|---|
| Managed WAF | 47.6% |
| CAPTCHA widget (reCAPTCHA, hCaptcha, Turnstile…) | 12.7% |
| Access control | 2.1% |
| Dedicated bot management | 0.8% |
| Any managed anti-bot / WAF | 53.5% |
This is not a fragmented market. Cloudflare alone covers 45% of every reachable site — 84% of all protected sites — and the next-largest named layer is consumer CAPTCHA widgets. The pure-play bot-management vendors barely register at web scale: Akamai Bot Manager 0.6%, DataDome 0.16%, PerimeterX/HUMAN 0.09%.
These numbers line up with independent measurements. W3Techs (June 2026) puts Cloudflare on 23.2% of all websites and roughly 48.7% of the top 1M — we find its active anti-bot posture on 45% of the reachable top 1M, the same ballpark from a different method. And the Imperva 2025 Bad Bot Report finds automated traffic is now 51% of all web traffic (bad bots 37%) — the demand pressure that explains why so many sites have a wall at all. (For a peer methodology on web-scale measurement, see the HTTP Archive Web Almanac.)
The busiest sites run the least anti-bot
The web-scale number hides the most useful finding. Split the million by traffic rank and three curves all move together — protection, Cloudflare and the specialist vendors:
| Tranco rank band | Protected | Cloudflare | Enterprise bot-mgmt¹ |
|---|---|---|---|
| Top 1,000 | 44.2% | 23.4% | 7.6% |
| 1k–10k | 50.7% | 34.0% | 5.3% |
| 10k–100k | 52.3% | 40.2% | 2.6% |
| 100k–1M | 53.6% | 45.6% | 0.6% |
¹ Akamai Bot Manager + DataDome + PerimeterX, as a share of reachable sites in the band.
Two things flip as you go down-rank. Cloudflare climbs (23% of the top-1,000 to 46% of the long tail): its web-scale dominance is a long-tail story — the giants run their own or enterprise infrastructure and keep homepages open for SEO, while the millions of smaller sites reach for turnkey Cloudflare. Meanwhile enterprise bot-management falls (7.6% to 0.6%): DataDome, PerimeterX and Akamai Bot Manager are a head phenomenon, bought by the high-value sites where automated data has a direct dollar cost. So the head and the tail aren't just protected at different rates — they're guarded by different vendors. (W3Techs sees the same direction: Cloudflare is lower in the very top-1,000 than across the broader top-1M.)
Set Cloudflare aside (it climbs across the board) and the specialist gatekeepers swap places by rank — Akamai guards the very top, while the long tail reaches for turnkey CAPTCHA:
Zoom into the ~1,005 high-value sites people actually scrape and the head's profile sharpens further: there, enterprise bot-management sits on 15.5% of sites, and the tail's cheap CAPTCHA-widget layer nearly vanishes. Adoption and sophistication are different axes: more of the web is "protected" than you'd think, but far less of it is hard.
Most of those walls are asleep
Here's the part that surprised us most. "53.5% protected" counts every site that exposes a managed vendor — but exposing one isn't the same as using it. Of the 818,614 reachable sites, only 79,835 (9.8%) actually challenged our homepage request; the other 358,022 walled sites ran the vendor passively, returning a clean 200 to a matched request. Follow the million all the way down to the few that actually fight back:
Show the flows
| Scanned → Reachable | 818,614 (36.3%) |
| Reachable → Has a wall | 437,857 (19.4%) |
| Reachable → No wall | 380,757 (16.9%) |
| Has a wall → Passively present | 358,022 (15.9%) |
| Scanned → Unreachable | 179,883 (8%) |
| Has a wall → Actively challenged | 79,835 (3.5%) |
Which wall is "awake" depends heavily on the vendor. Cloudflare sits in front of 45% of the reachable web but actively challenged only ~16% of those homepages; Google reCAPTCHA almost never fires on a homepage (3%); the small unidentified-WAF group is the opposite — present rarely, but challenging 76% of the time.
The practical takeaway: a vendor fingerprint tells you what you'd face if the wall fires — but most homepages won't fire at all, which is exactly why reaching for a full browser by default wastes time and budget. Check the page, not the logo.
Zoom in: the ~1,005 sites people actually scrape
For the sites that matter to a scraper, we go beyond headers: any site that blocks a datacenter GET is re-probed through the full transport fleet (browser-impersonation HTTP → headless browser → stealth browser + residential IP), so the tier reflects what actually reaches the page, not a header guess. By that bar, 575 of the 1,005 (57.2%) run managed bot defense, and two vendors dominate it:
| Vendor | Type | Sites | Share |
|---|---|---|---|
| Cloudflare | WAF + challenge | 333 | 33.1% |
| Akamai Bot Manager | Bot management | 111 | 11.0% |
| Akamai (edge) | CDN/WAF | 53 | 5.3% |
| DataDome | Bot management | 30 | 3.0% |
| Imperva (Incapsula) | WAF | 15 | 1.5% |
| PerimeterX (HUMAN) | Bot management | 15 | 1.5% |
| Cloudflare Turnstile | CAPTCHA | 12 | 1.2% |
Cloudflare and Akamai together account for 77% of every protected site in this set. What separates the vendors is what they inspect: Akamai and open Cloudflare paths lean hardest on TLS/JA3-JA4 fingerprinting — a matched HTTP client often reaches them — while DataDome and PerimeterX add real-time behavioral ML, so a clean fingerprint alone isn't enough.
Difficulty is wildly uneven — most of it is not a browser job
The more useful question isn't "is it protected" — it's "what does it take to get the public page reliably?" Running the real fleet against the curated set:
| Tier | What it takes | Sites | Share |
|---|---|---|---|
| T1 | Plain HTTP client | 608 | 60.5% |
| T2 | Browser-impersonation HTTP (matched TLS) | 246 | 24.5% |
| T3 | Headless browser that runs JavaScript | 148 | 14.7% |
| T4 | Stealth browser + residential IP + behavior | 3 | 0.3% |
85% of these sites need no browser at all — a plain HTTP GET or a matched TLS fingerprint reaches them. Only ~15% genuinely need a headless browser or more. Reaching for a browser by default is the most common (and most expensive) scraping mistake; the data says escalate only when a site forces you to.
This is also why "57% protected" and "85% need no browser" are both true: detecting a vendor is binary, but a vendor being present isn't the same as it actively challenging. Of the 575 protected sites, only 74 actively challenged our request — the rest run their vendor in passive CDN/WAF mode (a matched fingerprint returns 200).
The hard end: a few sites sign every request with a closed VM
The toughest class isn't a CAPTCHA — it's a proprietary in-browser bytecode VM that signs every request, so generic transport tooling can't mint a valid token. Four sites in the curated set ship one: TikTok's webmssdk VM (the X-Bogus / X-Gnarly signatures, on top of Akamai) and Kasada's proof-of-work VM (the x-kpsdk-ct token, on real-estate marketplaces). For these, "send a cleverer request" doesn't apply; you need a genuine browser execution context. They're rare, but they're the sites people most often ask "why can't I scrape this?" about. (We deliberately don't count an F5 BIG-IP load-balancer cookie as a VM — that's server load balancing, not bot defense.)
A block isn't a block — read why you were stopped
When a request doesn't pass, the reason tells you the fix — and the fix is completely different each time. After escalating through the fleet, only 74 of the 1,005 (7.4%) still didn't pass cleanly, and they break down as:
| Why it stopped | Sites | What it means | The fix |
|---|---|---|---|
| Bot challenge (JS interstitial) | 55 | Cloudflare "Just a moment", a JS challenge | Run it in a real browser |
| CAPTCHA | 18 | An interactive puzzle was served | A browser (+ CAPTCHA service) |
| Geo-blocked | 1 | Gated by country/region | An IP in an allowed region |
The headline is what the failures are not. Once you bring the right transport, genuine rate limits and outright IP bans are vanishingly rare at the front door — the wall is almost always a fingerprint/JS problem, not a proxy one. We were deliberately conservative on labels: a 401 only counts as a login wall if it carries WWW-Authenticate (WSJ and Reuters return 401 as a DataDome bot block, not authentication), and a generic 403 is "blocked," not "IP banned," unless the page says so.
The homepage is a poor proxy — protection moves per page
The biggest caveat in any study like this: a domain protects different page types completely differently. On LinkedIn, the homepage is light, a /company/ page is largely open, but a /in/ profile and search are login-walled. On Amazon, the homepage is open while search readily serves a "Robot Check." So every site page in the index ships an advisory deep-page test plan — which page types to check and what to expect on each — so you test the page you actually want, not the front door.
The defense gradient: crypto and commerce are fortresses, search is open
Split the curated set by category and a clear money gradient appears — the closer a page is to a transaction, the harder it's defended:
| Category | % managed anti-bot / WAF |
|---|---|
| Crypto | 86% |
| Marketplaces | 81% |
| AI tools | 78% |
| Finance & markets | 75% |
| Forums & community | 75% |
| E-commerce | 72% |
| Real estate | 72% |
| Travel & hospitality | 72% |
| … | |
| News & media | 39% |
| Sports | 33% |
| Social media | 25% |
| Search engines | 18% |
Crypto and marketplaces fight hardest — prices, inventory and order books are the most-scraped data on the web. The low numbers for social, news and search are partly a method artifact: those homepages are deliberately open, but the deep pages (a profile, an article archive, a result page) flip to protected the moment you navigate in.
What this means if you're building a scraper
- Check before you build. Knowing a page is behind Cloudflare versus DataDome versus a closed VM versus nothing changes the entire approach — and it changes per URL. A 30-second check saves an afternoon of wondering why your requests
403. See the companion guide on scraping sites that block bots for what each vendor actually does. - Escalate only as far as the site demands, and pay only when it works. 85% of the sites people scrape don't need a browser at all. The cost-efficient pattern is to try the lightest transport first and escalate, which is exactly how Crawlora's web scraping for AI and data pipelines is billed — pay-on-success, not per attempt.
None of this is about defeating a CAPTCHA or getting past a login. It is about knowing, before you commit engineering time, which public pages need which approach.
Try it yourself
- Check any URL right now with the free anti-bot checker — paste a profile or listing page, not just the homepage, and see the vendor and the lightest transport that works.
- Explore the full dataset — search every site, filter by difficulty, vendor or category — in the Anti-Bot Adoption Index.
- Read exactly how we measured it, including the signatures and the difficulty heuristic.
How the check actually works (full transparency)
No magic — this is a deliberately simple, passive, reproducible check, run at two scales.
The samples. The web-scale figures come from the full Tranco top 1,000,000 (an aggregated, citable top-sites list with a permanent ID), of which 818,614 were reachable. The deep-dive figures come from 1,005 domains walked from the top of the same ranking into real, public, content-bearing sites across 28 categories, with pure infrastructure/CDN, ad/tracking and adult domains skipped.
The request. Each site's homepage is fetched with a real Chrome User-Agent, following redirects, from a datacenter IP — the honest "what a basic cloud scraper sees" vantage. We never submit a form, run a CAPTCHA, log in, or fetch anything behind a wall. The top-1M scan is this single datacenter GET. For the curated set we go further: any site that blocks the datacenter GET is re-probed through the full transport fleet (browser-impersonation → headless → stealth + residential), so its difficulty tier reflects what actually reaches the page rather than a header guess.
What we capture (and what we don't). From the response we keep the status code, the response headers, the names of any Set-Cookie cookies — names only, never values, so no session tokens or PII are stored — and a capped slice of the body. Nothing else is retained.
Naming the vendor. The vendor is identified by matching that evidence against a database of public, documented fingerprints — header names (cf-ray, x-datadome, x-iinfo, akamai-grn), Set-Cookie name prefixes (__cf_bm/cf_clearance, _abck/bm_sz, datadome, _px*, incap_ses_), and body markers that are only trusted on a challenge-shaped response. Header and cookie matches are high-confidence; body markers medium.
Limits. Every result is a lower bound — homepages are more open than the deep pages people actually scrape, and a datacenter IP sees more challenges than a residential one. The top-1M scan is headers-only and does not grade difficulty (a site that blocks a datacenter GET there may well open for a browser); the graded tiers and the category gradient come from the full-fleet curated run. Anti-bot deployments change continuously, so treat every result as a directional signal, not a guarantee. Snapshot: June 2026.
Open & reproducible. The classifier and the methodology are published in full — read the full methodology, or explore the searchable index. Cite it as "Crawlora Anti-Bot Adoption Index" with a link.
Frequently asked questions
How much of the web is protected by anti-bot?
We measured it twice. Across the full Tranco top 1,000,000 (818,614 reachable), 53.5% expose a managed anti-bot, WAF or access-control vendor — overwhelmingly Cloudflare (45% of reachable sites). On 1,005 hand-picked high-value sites run through the full transport fleet, 57.2% are protected. Both are lower bounds: deep pages (profiles, listings, search) are more defended than the homepages we probed, and protection is heaviest on crypto (86%) and marketplaces (81%), lightest on search, social and news.
Do you actually need a browser to scrape most sites?
No — and that's the most expensive misconception. Of the curated sites, 85% need no browser at all: about 60% answer a plain HTTP GET and another 25% only want a matched TLS fingerprint. Only ~15% genuinely need a headless browser or more. The cost-efficient pattern is to escalate only as far as a site forces you to, not to reach for a headless browser by default.
Which anti-bot vendor is the most common?
Cloudflare, by a wide margin — across the top 1,000,000 it covers 45% of reachable sites and 84% of every protected site. In the curated set it leads at 33%, followed by Akamai (Bot Manager 11% plus its edge CDN 5%). The specialist bot-management vendors DataDome (3%) and PerimeterX/HUMAN (1.5%) cluster on high-value verticals like marketplaces, travel and real estate, where automated data has a direct cost.
What is the hardest type of anti-bot to deal with?
A proprietary, closed in-browser JavaScript VM that signs every request — like TikTok's webmssdk (the X-Bogus/X-Gnarly signatures) or Kasada (the x-kpsdk-ct token); four sites in the curated set ship one. Unlike a CAPTCHA, generic transport tooling can't mint a valid token, so these need a genuine browser execution context rather than a cleverer request. They're rare but they're the sites people most often ask why they can't scrape.
How did you measure each site's anti-bot?
Two ways. The top-1,000,000 scan is a single passive homepage GET from a datacenter IP — matching response headers (cf-ray, x-datadome), Set-Cookie names (_abck, datadome, _px) and challenge markers against documented vendor fingerprints, the same signatures as Crawlora's anti-bot checker. For the curated set we go further: any site that blocks the datacenter GET is re-probed through the full transport fleet (browser-impersonation → headless → stealth + residential), so its difficulty tier reflects what actually reaches the page, not a header guess. Every result is a lower bound.