Cloudflare Blocks AI Crawlers by Default: A Publisher Guide to Content Controls
Cloudflare now blocks AI crawlers by default for new domains and is piloting Pay Per Crawl. With IAB Tech Lab and Google also moving, here is how publishers should set an AI access policy.
For two years, publishers have watched AI systems consume their content while sending back a trickle of traffic. In the past five weeks, three moves gave them new tools to respond. On July 1, Cloudflare announced that it will block AI crawlers by default, asking new domains up front whether to allow them, and launched a private beta of Pay Per Crawl, which lets site owners charge AI crawlers for access. In June, the IAB Tech Lab announced an LLM Content Ingest API initiative aimed at standards for access control and compensation. And on June 26, Google launched Offerwall in Google Ad Manager, a tool that gives readers several ways to access content, including watching an ad or making a micropayment.
None of these solves the problem on its own. Together they mark a shift: publishers now have practical ways to decide who can use their content and on what terms. The question is what policy to set.
Why the numbers pushed Cloudflare to act
Cloudflare's announcement argued that the exchange between crawling and traffic has broken down. It said that, by its measurements, getting traffic from OpenAI was 750 times harder, and from Anthropic 30,000 times harder, than it was from the Google of old. Those are Cloudflare's figures and its framing, but they match what many publishers see in their logs: heavy crawler activity, very few referrals.
With a large share of the web running through Cloudflare, making blocking the default for new domains changes the starting point of the negotiation. AI companies that want the content increasingly need permission.
What each new tool does
Cloudflare: default blocking and Pay Per Crawl
Default blocking means AI crawlers that Cloudflare identifies are stopped unless the site owner allows them. Pay Per Crawl, in private beta, lets a site set a price for access; crawlers that do not pay receive an HTTP 402 Payment Required response. Existing Cloudflare customers can adjust settings for their domains.
IAB Tech Lab: standards for access and payment
The Tech Lab's proposal, announced at its June summit, covers content access controls using robots.txt and firewall rules, ways for AI systems to discover publisher content and terms, and monetization models including cost per crawl and an ingest API. It is a framework under development, not a finished standard, and the Tech Lab has invited publishers and AI companies to participate.
Google Offerwall: monetizing readers who would otherwise leave
Offerwall addresses a different piece of the problem: the value of each human visit. It lets publishers offer readers choices, such as watching a short ad, paying a small amount through Supertab, answering a survey or signing up for a newsletter, in exchange for access. Google said it tested Offerwall with about 1,000 publishers for more than a year and reported an average revenue uplift of 9%. It is available in Ad Manager at no cost.
Setting an AI access policy
Before changing any settings, decide what you want. A simple framework:
- Inventory the bots. Use server or CDN logs to list which AI crawlers visit, how often, and which sections they hit. Distinguish training crawlers, search and retrieval crawlers, and user-triggered agents.
- Separate search from training. Blocking a crawler that powers an AI product's live answers may reduce visibility there. Blocking training crawlers has less immediate traffic impact. Treat them differently.
- Keep search engines working. Do not block Googlebot or Bingbot in an attempt to stop AI features. That removes you from search results. Use crawler-specific tokens, such as Google-Extended for Gemini model training, where they exist.
- Decide where licensing fits. If you have content AI companies value, such as archives, niche expertise or data, blocking by default gives you leverage for a licensing conversation. If you do not plan to negotiate, blocking still protects bandwidth and origin costs.
- Write it down. Publish your terms in robots.txt and, where relevant, on a content access page. Clear terms make enforcement and future deals easier.
A note on AI agents
Crawlers that collect content for training are only one kind of automated visitor. AI assistants and browsing agents increasingly fetch pages on behalf of a specific user who asked a question. These visits look different in logs, may not load ads at all, and raise their own questions about attribution and payment. For now, the practical step is visibility: identify these agents in your logs, measure how often they visit and whether they render ads, and keep them in mind when you set policy. The standards being developed by the IAB Tech Lab and others are likely to address them more directly over time.
Robots.txt is not enough on its own
Robots.txt is a request, not a lock. Well-behaved crawlers honor it; others ignore it or disguise themselves. That is why network-level enforcement, whether through a CDN, a web application firewall or bot management, matters. If you are not on Cloudflare, check what your CDN offers for verified bot detection and AI crawler rules.
Do not forget the human visitors
Controlling crawlers protects your content. It does not bring back lost traffic. The other half of the strategy is making every human visit more valuable: better session depth, well-placed ads, value exchanges like Offerwall or rewarded formats for readers who hit a paywall, and newsletters that bring readers back without an intermediary.
What to do this month
- Pull 30 days of logs and quantify AI crawler traffic by bot.
- Decide your policy for training, search and agent crawlers separately.
- Update robots.txt and CDN or firewall rules to match.
- Evaluate value-exchange tools for non-subscribers.
- Follow the IAB Tech Lab framework as it develops.
The balance of power between publishers and AI platforms is not settled, but the default is starting to shift. Publishers who set a clear policy now will be in a better position whatever the market for content access becomes. HBDR focuses on the monetization side of that equation, making each human visit count, while publishers decide how their content is licensed.
Related Articles
AAMP 3.0 and OpenProposal: Getting Your Inventory Ready for Agent-Written RFPs
IAB Tech Lab's AAMP 3.0 introduces OpenProposal, a standard way for buyer and seller agents to exchange briefs and proposals. What it means for publishers who sell directly.
Cloudflare's Sept. 15 AI Crawler Deadline: What Recipe Publishers Should Check
Starting September 15, Cloudflare's defaults block mixed-use AI crawlers from pages that host ads. What that means for recipe and lifestyle publishers, and the settings worth reviewing.
No Breakup for Google Ad Tech: What the Remedies Ruling Means for Publishers
Judge Brinkema rejected a forced sale of AdX and ordered behavioral remedies instead. What the September 2 ruling changes for publishers, and what to do before it takes effect.