Back to Blog
Industry Trends September 15, 2025 4 min read

RSL, the Penske Lawsuit and the New Economics of AI Crawlers

A new licensing standard for AI crawlers and the first major U.S. publisher lawsuit over AI Overviews landed in the same week. Here is what publishers can control today.

HR
HBDR Research
September 15, 2025

Two moves in one week

Publishers have spent most of 2025 watching AI answers absorb clicks that used to come from search. Last week produced two responses that point in different directions: one technical and cooperative, one legal and adversarial. Both are worth understanding, because together they frame the choices publishers will make about AI crawlers for the next few years.

RSL: a licensing layer for the web

On September 10, a group of publishers and technology companies launched Really Simple Licensing (RSL), an open standard that lets sites publish machine-readable licensing terms for AI systems. It is managed by the nonprofit RSL Collective, co-founded by RSS co-creator Eckart Walther and former Ask.com CEO Doug Leeds. Launch supporters included Reddit, Yahoo, Medium, Quora, O'Reilly Media, Ziff Davis, People Inc., Internet Brands and Fastly.

RSL lets a publisher declare terms, primarily through robots.txt, that range from free use, to attribution, to pay-per-crawl, to pay-per-inference, where compensation is tied to each time an AI system uses the content in an answer. The idea is to give AI companies a clear way to license content and give publishers a standard way to state their terms, rather than negotiating one deal at a time.

The obvious limitation is enforcement. A license file is a statement of terms, not a lock. It works if crawlers read and respect it, or if infrastructure providers enforce it on the publisher's behalf.

Enforcement is already moving to the network layer

That is why the infrastructure side matters. On July 1, Cloudflare began asking new domains whether to allow AI crawlers, making blocking the default for new sign-ups, and it launched a private beta of pay per crawl, which uses the HTTP 402 Payment Required status to let sites charge crawlers for access.

In August, Cloudflare published a report alleging that Perplexity used undeclared crawlers, with a generic browser user agent and rotating IPs, to reach sites that had blocked its declared bots. Cloudflare estimated 20–25 million daily requests from the declared crawler and 3–6 million from the undeclared one. Perplexity disputed the findings. Whatever the merits, the episode showed that robots.txt alone is not a reliable control, and that detection and blocking increasingly happen at the CDN and bot-management layer.

The Penske lawsuit

On September 12, Penske Media Corporation, which owns Rolling Stone, Billboard, Variety and other titles, sued Google in federal court in Washington, D.C. It is the first major U.S. media company to challenge Google over AI Overviews. The core allegation is that Google uses its search monopoly to require publishers to supply content for AI summaries as a condition of being included in search results, which reduces the clicks publishers depend on. According to reporting on the complaint, about 20% of Google searches that link to a Penske site now show AI Overviews, and Penske says its affiliate revenue fell by more than a third from its peak by the end of 2024. Google has said AI Overviews send traffic to a broader range of sites.

The structural point in the complaint is one many publishers already feel: Google offers a Google-Extended control that governs whether content is used to train Gemini models, and Google states it does not affect inclusion or ranking in Search. But there is no equivalent switch that lets a site stay in regular search results while opting out of AI Overviews without also limiting how its pages appear in search. That is the leverage the lawsuit targets.

What publishers can control today

Litigation will take years. Standards will take time to be adopted. Meanwhile, there are concrete steps worth taking.

  1. Measure crawler load. Pull a week of server or CDN logs and group requests by declared bot. Many publishers are surprised by how much bandwidth AI crawlers consume relative to the referral traffic they send back.
  2. Set a policy by crawler purpose. Search indexing, model training and real-time retrieval for answers are different uses with different value to you. Decide which you allow, block or would license, and express that in robots.txt and, if you choose, RSL terms.
  3. Enforce at the edge. If you use a CDN with bot management, use it. Verified-bot lists and behavior-based detection catch what robots.txt cannot.
  4. Track referral quality, not just volume. Sessions arriving from AI products may be fewer but deeper. Measure pages per session and revenue per session by referrer.
  5. Keep your ad stack out of bot traffic. Undeclared crawlers that render pages can trigger ad requests. Make sure your invalid traffic filtering and verification partners are catching them, so buyers do not discount your inventory.

For entertainment and lifestyle publishers

Entertainment, music and lifestyle sites are especially exposed, because many of their search visits are short informational queries that an AI summary can answer on the page. For these publishers the priority is to grow sessions that do not start from search: newsletters, social, direct and app traffic. Those users also tend to be the most valuable to advertisers.

The takeaway

The industry is testing three levers at once: licensing standards, network-level enforcement and the courts. None of them will restore lost search traffic on its own this year. What publishers can do is know exactly who is crawling them, decide deliberately what to allow, and make each remaining pageview earn as much as possible. A well-tuned ad stack, whether run in-house or through a managed partner such as HBDR, is part of that last step.

Tags: ai search content licensing crawlers rsl publisher traffic

Ready to maximize your ad revenue?

Get Started