Tools

Cloudflare lets publishers separate search, training and agent bot access; monetization and BotBase introduced

Cloudflare has rolled out new controls allowing web publishers to manage crawler access by purpose—search indexing, AI training, or agent activity—and will block training and agent bots by default while permitting search crawlers.

Cloudflare lets publishers separate search, training and agent bot access; monetization and BotBase introduced

Web publishers using Cloudflare will soon be able to control crawler access by purpose: separately allowing search indexing, blocking AI-training crawlers, or limiting agent bots that act on users’ behalf. Cloudflare also announced a public bot registry and a pay-per-request monetization feature that is currently available by waitlist.

What’s changing?

  • Cloudflare, which serves roughly 20 percent of the web, unveiled new AI traffic controls. By default, crawlers used for AI training and agent activity will be blocked, while search-indexing crawlers will remain permitted.
  • The company also introduced a monetization mechanism meant to allow customers to charge specified visitors—human users or bots—for access to web pages, datasets, APIs, or Model Context Protocols. That monetization tool is not yet released; customers can join a waitlist to use it.

How the controls work

Cloudflare customers can set access per crawler use case (search, AI training, AI agents) with three options for each category: block all pages, block only on pages that show ads, or allow the crawler. Crawlers that serve multiple purposes will be governed by the most restrictive choice. For example, if a publisher blocks training crawlers, bots that both index content for search and collect training data—such as Googlebot, Applebot, and Bingbot—would be blocked.

The controls will be applied starting September 15, 2025.

Cloudflare defines search bots as those that index content for search engines, agent bots as systems that perform tasks on a site on a user's behalf (for instance fetching data or completing purchases), and training bots as those that collect data to train AI models.

BotBase and the “Verified” label

Cloudflare also launched BotBase, a public database of known bots and agents that classifies them by behavior and tracks how they use website content. A bot labeled “Verified” indicates the operator has been authenticated, the bot corresponds to its declared identity, obeys robots.txt, and does not try to circumvent publisher restrictions. Publishers still decide whether to allow a given category (search, training, or agents), but BotBase enables more granular policies—such as blocking search, training, or agent activity from a particular company while allowing others.

Monetization at the network level

The new monetization system would let publishers charge AI agents for per-request access to online data. When an AI agent requests a protected page or API, Cloudflare would require payment verification before forwarding the request to the publisher, handling that at the network layer. The feature is not yet live; interested customers can join a waitlist.

Why this matters

Cloudflare reports that automated systems now generate the majority of HTTP requests—57.5 percent—on the web. Some publishers say the value of search traffic has declined, and a few have opted out of Google Search in part to prevent search agents from bypassing display advertising or using their content for model training.

Cloudflare’s framework seeks to preserve the benefits of search-driven traffic while giving publishers control over other forms of AI crawling. The approach advantages companies, such as OpenAI, that separate their search and training crawlers and penalizes those that do not, including Google, Apple, and Microsoft.

Potential consequences and debate

Access to high-quality web data remains central for developing AI models, but increasing limits on data collection introduce technical and monetary costs. Those costs could disproportionately affect smaller and newer AI developers who lack the resources to negotiate publisher agreements or pay for access, compared with larger firms that have already integrated extensive web data into their models.

Cloudflare positions itself as a potential intermediary—a toll keeper—between AI companies and publishers, but the change challenges the widely held AI-industry view that training on public web data is fair use. It remains uncertain whether Cloudflare has the leverage to enforce such charges and which interpretation of “fairness” will ultimately prevail.