Back to Blog
Industry Trends July 7, 2025 5 min read

Cloudflare Blocks AI Crawlers by Default: A Publisher Guide to Content Controls

Cloudflare now blocks AI crawlers by default for new domains and is piloting Pay Per Crawl. With IAB Tech Lab and Google also moving, here is how publishers should set an AI access policy.

HR
HBDR Research
July 7, 2025

For two years, publishers have watched AI systems consume their content while sending back a trickle of traffic. In the past five weeks, three moves gave them new tools to respond. On July 1, Cloudflare announced that it will block AI crawlers by default, asking new domains up front whether to allow them, and launched a private beta of Pay Per Crawl, which lets site owners charge AI crawlers for access. In June, the IAB Tech Lab announced an LLM Content Ingest API initiative aimed at standards for access control and compensation. And on June 26, Google launched Offerwall in Google Ad Manager, a tool that gives readers several ways to access content, including watching an ad or making a micropayment.

None of these solves the problem on its own. Together they mark a shift: publishers now have practical ways to decide who can use their content and on what terms. The question is what policy to set.

Why the numbers pushed Cloudflare to act

Cloudflare's announcement argued that the exchange between crawling and traffic has broken down. It said that, by its measurements, getting traffic from OpenAI was 750 times harder, and from Anthropic 30,000 times harder, than it was from the Google of old. Those are Cloudflare's figures and its framing, but they match what many publishers see in their logs: heavy crawler activity, very few referrals.

With a large share of the web running through Cloudflare, making blocking the default for new domains changes the starting point of the negotiation. AI companies that want the content increasingly need permission.

What each new tool does

Cloudflare: default blocking and Pay Per Crawl

Default blocking means AI crawlers that Cloudflare identifies are stopped unless the site owner allows them. Pay Per Crawl, in private beta, lets a site set a price for access; crawlers that do not pay receive an HTTP 402 Payment Required response. Existing Cloudflare customers can adjust settings for their domains.

IAB Tech Lab: standards for access and payment

The Tech Lab's proposal, announced at its June summit, covers content access controls using robots.txt and firewall rules, ways for AI systems to discover publisher content and terms, and monetization models including cost per crawl and an ingest API. It is a framework under development, not a finished standard, and the Tech Lab has invited publishers and AI companies to participate.

Google Offerwall: monetizing readers who would otherwise leave

Offerwall addresses a different piece of the problem: the value of each human visit. It lets publishers offer readers choices, such as watching a short ad, paying a small amount through Supertab, answering a survey or signing up for a newsletter, in exchange for access. Google said it tested Offerwall with about 1,000 publishers for more than a year and reported an average revenue uplift of 9%. It is available in Ad Manager at no cost.

Setting an AI access policy

Before changing any settings, decide what you want. A simple framework:

  1. Inventory the bots. Use server or CDN logs to list which AI crawlers visit, how often, and which sections they hit. Distinguish training crawlers, search and retrieval crawlers, and user-triggered agents.
  2. Separate search from training. Blocking a crawler that powers an AI product's live answers may reduce visibility there. Blocking training crawlers has less immediate traffic impact. Treat them differently.
  3. Keep search engines working. Do not block Googlebot or Bingbot in an attempt to stop AI features. That removes you from search results. Use crawler-specific tokens, such as Google-Extended for Gemini model training, where they exist.
  4. Decide where licensing fits. If you have content AI companies value, such as archives, niche expertise or data, blocking by default gives you leverage for a licensing conversation. If you do not plan to negotiate, blocking still protects bandwidth and origin costs.
  5. Write it down. Publish your terms in robots.txt and, where relevant, on a content access page. Clear terms make enforcement and future deals easier.

A note on AI agents

Crawlers that collect content for training are only one kind of automated visitor. AI assistants and browsing agents increasingly fetch pages on behalf of a specific user who asked a question. These visits look different in logs, may not load ads at all, and raise their own questions about attribution and payment. For now, the practical step is visibility: identify these agents in your logs, measure how often they visit and whether they render ads, and keep them in mind when you set policy. The standards being developed by the IAB Tech Lab and others are likely to address them more directly over time.

Robots.txt is not enough on its own

Robots.txt is a request, not a lock. Well-behaved crawlers honor it; others ignore it or disguise themselves. That is why network-level enforcement, whether through a CDN, a web application firewall or bot management, matters. If you are not on Cloudflare, check what your CDN offers for verified bot detection and AI crawler rules.

Do not forget the human visitors

Controlling crawlers protects your content. It does not bring back lost traffic. The other half of the strategy is making every human visit more valuable: better session depth, well-placed ads, value exchanges like Offerwall or rewarded formats for readers who hit a paywall, and newsletters that bring readers back without an intermediary.

What to do this month

  • Pull 30 days of logs and quantify AI crawler traffic by bot.
  • Decide your policy for training, search and agent crawlers separately.
  • Update robots.txt and CDN or firewall rules to match.
  • Evaluate value-exchange tools for non-subscribers.
  • Follow the IAB Tech Lab framework as it develops.

The balance of power between publishers and AI platforms is not settled, but the default is starting to shift. Publishers who set a clear policy now will be in a better position whatever the market for content access becomes. HBDR focuses on the monetization side of that equation, making each human visit count, while publishers decide how their content is licensed.

Tags: ai crawlers cloudflare content licensing robots.txt offerwall

Ready to maximize your ad revenue?

Get Started