The State of Web Scraping & Anti-Bot 2026
We scanned the top 1 to 10 million sites to map how the open web defends itself: 53.5% run an anti-bot wall, 9.33% block AI crawlers, and 14.1% are dead.
Every year a new report tells you that bots are now most of web traffic. That number is real, but it only describes one side of the relationship. We wanted the other side: not how much of the traffic is automated, but how much of the web itself is now built to keep automated clients out.
So we ran three population-scale scans in June 2026 — the Anti-Bot Adoption Index, the AI-Crawler Blocking Index, and the Dead-Web Index — across the top 1 to 10 million sites. Read together, they describe a web that is getting measurably harder to read by machine. Here are the three findings; the full report has the charts, the methodology, and the open data.
1. The web is walling up
Of 998,497 sites in the Tranco top 1M, 818,614 were reachable, and 53.5% of those run a managed anti-bot or WAF wall. Cloudflare alone is about 45% of the reachable web and 84% of every protected site.
The counterintuitive part is the gradient. The most popular sites are the least likely to run a commodity wall — the top 1,000 sit at 44%, the long tail at 54%. The giants are not unprotected; they run bespoke enterprise bot-management instead (DataDome, Akamai, PerimeterX), which fades from ~7.6% at the head to 0.6% at the tail. Commodity Cloudflare adoption climbs the other way.
| Traffic rank | Run a managed wall | Cloudflare | Enterprise bot-mgmt |
|---|---|---|---|
| Top 1K | 44.2% | 23.4% | 7.58% |
| 1K–10K | 50.7% | 34.0% | 5.29% |
| 10K–100K | 52.3% | 40.2% | 2.55% |
| 100K–1M | 53.6% | 45.6% | 0.60% |
By extension the spread is even wider: the Anglophone web is the most walled (.us 89%), the East-Asian web the most open (.kr 11%) — an eightfold range.
2. AI crawlers are gated selectively
Despite the headlines about publishers blocking AI, only 9.33% of the top 1M fully block even one major AI crawler in robots.txt (14.8% of the sites that publish a robots.txt at all). And the blocking is deliberate: sites shut out the crawlers that train models far more than the agents that send readers back.
OpenAI's GPTBot leads at 7.4%, but Common Crawl's CCBot — the open corpus most models train on indirectly — is a near-tie at 7.23%. The AI-search and assistant agents that drive referral traffic are barely blocked.
| AI crawler | Purpose | Sites fully blocking |
|---|---|---|
| GPTBot (OpenAI) | Training | 7.40% |
| CCBot (Common Crawl) | Training | 7.23% |
| Bytespider (ByteDance) | Training | 6.77% |
| ClaudeBot (Anthropic) | Training | 6.69% |
| Google-Extended | Training | 6.28% |
| ChatGPT-User (OpenAI) | AI assistant | 1.59% |
| PerplexityBot | AI search | 1.29% |
Unlike anti-bot walls, AI-blocking is concentrated at the head — the busiest, most content-heavy sites — and led by news & media, which block at roughly 81%.
3. The dead web is overstated
You have probably seen the stat that a quarter of the web is gone. We probed all 10 million most-popular domains (9.99M reached, 99.95% coverage) and get 14.1% genuinely dead — about half the usual figure.
The gap is method. A naive crawler counts a 403, a 429, or a served 404 as death, but those sites are alive and merely blocking bots (8.9% answer but block), or misconfigured. True death is DNS: 76% of dead domains no longer resolve at all. And it is a tail phenomenon — 99.8% of dead domains sit below rank 100,000, so weighted by traffic the dead web is closer to 3%.
| Outcome | Share | What it means |
|---|---|---|
| Alive | 76.6% | responds with usable content |
| Dead | 14.1% | DNS gone or no server (76% no longer resolve) |
| Blocked | 8.9% | answers but blocks bots (403 / 429) |
| Redirect | 0.3% | redirects off the original domain |
What it means
The open web is not dying. It is professionalizing its defenses, selectively gating AI, and decaying at the tail. For anyone who builds on public web data, the practical lesson is the same one all three scans point to: reliable collection now depends on handling the walls, not just fetching pages — because more than half the web is now behind one.
The full report has every chart and the methodology, and all three datasets are open under CC BY 4.0 via the Anti-Bot, AI-Crawler, and Dead-Web indexes.
Frequently asked questions
How much of the web runs anti-bot protection in 2026?
In an independent June 2026 scan of the Tranco top 1 million sites (998,497 scanned, 818,614 reachable), 53.5% of reachable sites run a managed anti-bot or WAF wall, and Cloudflare alone is about 84% of every protected site. Protection rises as you go down the ranking: 44% of the top 1,000 versus 54% of the 100K to 1M band, because the largest sites use enterprise bot-management instead of a commodity wall.
Which AI crawlers do websites block the most?
GPTBot (OpenAI) is the single most-blocked at 7.4% of all top-1M sites, with Common Crawl's CCBot a near-tie at 7.23%. Overall only 9.33% of sites block at least one major AI crawler. Training crawlers are blocked roughly 5 times more than the AI-search and assistant agents (PerplexityBot, ChatGPT-User, OAI-SearchBot) that send referral traffic back.
How much of the web is actually dead?
14.1% of the top 10 million domains are genuinely dead, far below the often-cited 27.6% from naive crawls. The difference is method: a 403, a 429, or a served 404 is alive-but-blocking (8.9% answer but block bots) or misconfigured, not dead. True death is overwhelmingly DNS failure: 76% of dead domains no longer resolve at all.
How is this different from the Imperva, HUMAN, or Cloudflare bot reports?
Those reports measure bot traffic volume, the share of requests that come from bots (around 53% of all traffic). This report measures the opposite side of the ledger: the web's defensive posture and reachability, site by site, across the top 1 to 10 million domains. It answers how much of the web is walled, AI-gated, or gone, rather than how much traffic is automated.
