AI Crawlers Hit 50 Billion Requests a Day at Cloudflare. Check Your Ad Traffic
Cloudflare's new AI Labyrinth targets bots that ignore no-crawl rules. For publishers, crawler traffic is also an ad quality issue: inflated page views, wasted auctions and IVT risk.
What Cloudflare announced
On March 19, 2025, Cloudflare introduced AI Labyrinth, an opt-in feature that responds to unauthorized crawling by sending bots into a maze of AI-generated decoy pages rather than simply blocking them. The pages are linked in ways humans don't see, carry directives so search engines do not index them, and serve as a honeypot: anything that follows those links is a bot. It is available to all Cloudflare customers, including those on the free plan. In its announcement, Cloudflare said AI crawlers generate more than 50 billion requests to its network every day, just under 1% of all web requests it sees.
That number is a reminder that a meaningful share of what hits publisher servers is not people. For ad-supported sites, that is more than a bandwidth problem.
Why crawlers are an ad operations issue
Most AI crawlers fetch HTML and never execute JavaScript, so they do not trigger ad requests. But not all automated traffic is so polite, and the gray zone creates several problems:
- Distorted metrics. Bot page views inflate traffic in some analytics setups and depress page RPM, viewability and engagement rates, making real performance harder to read.
- Wasted auctions. Headless browsers that execute your pages run header bidding and ad requests. Bidders spend resources on impressions no person sees, and if they detect it, they may bid less on your inventory as a whole.
- Invalid traffic flags. Industry standards define general invalid traffic (GIVT) to include known crawlers and data center traffic, which ad servers and verification vendors filter. Sophisticated invalid traffic that mimics humans is harder to catch and more damaging if attributed to your site. A sustained spike can trigger partner reviews or clawbacks.
- Server cost and speed. Heavy crawling increases hosting costs and can slow pages for real users, which feeds back into Core Web Vitals and ad performance.
Which sites feel it most
Content that is useful for training and for answering questions gets crawled hardest. Recipe and lifestyle sites, with large archives of structured, evergreen content, are prime targets. So are reference, how-to, health information and B2B knowledge bases. These are also verticals where ad revenue per page is the core business, so the stakes of polluted traffic are higher.
A practical response
1. Measure first
Pull server or CDN logs for a recent week and group requests by user agent and verified source. Most major AI companies publish their crawler names, and many publish IP ranges. Compare total requests to the page views in your analytics and to ad requests in your ad server. Large gaps between these numbers are worth explaining.
2. Decide your policy by crawler, not all at once
Not every automated visitor is the same. Search crawlers bring referral traffic. Some AI assistants fetch pages in real time to answer a user's question and may link back. Pure training crawlers take content without sending visitors. Many publishers now distinguish between them in robots.txt, using the published tokens for crawlers such as GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended. Note that robots.txt is a request, not an enforcement mechanism; bots that ignore it are exactly what tools like AI Labyrinth and bot management products target.
3. Enforce at the edge
If you use a CDN with bot management, use it. Cloudflare has offered a one-click setting to block AI scrapers and crawlers since mid-2024, and other CDNs and security vendors offer similar controls. Edge enforcement stops unwanted bots before they reach your origin, your analytics or your ad stack.
4. Keep bots out of the auction
For automation that does get through, add safeguards on the ad side:
- Delay auctions until basic signals of a real visit, such as the page becoming visible, rather than firing on raw page load in hidden or background contexts.
- Use lazy loading so below-the-fold slots only request ads when a viewport actually approaches them.
- Filter known data center traffic from ad calls where your stack supports it, and review your ad server's invalid traffic reports monthly.
- Ask your SSPs for IVT rates by domain and investigate any upward trend quickly.
5. Clean up your reporting
Make sure analytics bot filtering is enabled and that dashboards for revenue per session and RPM are based on filtered human traffic. Decisions about layout, floors and partners should never be made on bot-inflated data.
Questions to ask your partners
- Does your SSP filter data center and known-bot traffic before sending bid requests, and can it report filtered volume for your domains?
- Does your analytics platform exclude known bots by default, and is that setting on for every property?
- Does your CDN or security vendor classify AI crawlers separately from search crawlers, so you can apply different rules to each?
The answers determine whether a crawler surge shows up as a clean line in a log or as a mystery drop in RPM weeks later.
The bigger picture
The questions of who can crawl publisher content, for what purpose and on what terms are moving quickly, from infrastructure providers building controls to licensing deals between publishers and AI companies. Whatever position your business takes on AI access, the ad operations principle is the same: auctions and impressions should be reserved for people.
Every auction run for a bot is noise in your data and a reason for buyers to trust your inventory a little less.
Start with the logs. If a managed partner runs your stack, ask for a joint review of IVT rates and bot filtering before the traffic mix shifts further.
Related Articles
July's Privacy Law Changes: A Checklist for Health and Finance Publishers
Connecticut, Arkansas and Virginia changes took effect July 1, and IAB Tech Lab just proposed GPP updates. What health, finance and other sensitive-content publishers should check.
Privacy Sandbox Is Winding Down in Chrome. Time to Clean Up Your Wrapper
Chrome 150 is now rejecting Protected Audience calls as Google retires most Privacy Sandbox APIs. Here is what to remove from your ad stack and what stays the same.
July 1 Privacy Deadlines: Connecticut and Arkansas Tighten Teen Ad Rules
On July 1, Connecticut's amended privacy law and Arkansas's children's and teens' privacy law take effect, both restricting targeted ads to minors. What publishers should change first.