Starting today, September 15, Cloudflare begins enforcing, by default, a block on "mixed-use" AI crawlers on any page that displays ads — a change announced in July that now officially takes effect for new domains on the platform, new sites from existing customers, and the entire free-tier base.

The "three-in-one" crawler problem

According to Cloudflare, a large share of today's AI bot traffic comes from crawlers that combine three jobs under a single user-agent: indexing for traditional search, fetching pages for real-time AI agents, and collecting data to train models. That mix made it nearly impossible for a site to keep showing up on Google while blocking the same content from training a competing model — because, until now, it was the same bot doing both.

As of today, Cloudflare splits traffic into three categories — Search, Agent and Training — and, for affected accounts, only the Search category remains allowed by default on ad-supported pages; Agent and Training are now blocked unless the site explicitly opts in.

From "pay to access" to "pay for use"

The change also advances the second stage of a plan the company has been building since 2025: the Pay Per Crawl program, which charged AI companies per page fetched, is evolving into a Pay Per Use model, where payment to the site happens when the content is actually used inside an AI answer — not simply when it's fetched. Cloudflare says more than half of today's AI crawler traffic is spent re-fetching pages that haven't changed, which made the per-fetch billing model inefficient for both publishers and AI companies. The new Pay Per Use marketplace already counts two initial partners, Ceramic.ai and You.com.

Why it matters for AI builders

For teams building agents or search pipelines that depend on accessing third-party content, today's change is a practical reminder: crawlers still using a single user-agent to fetch content for multiple purposes will now be blocked by default across a meaningful slice of the web. Clearly separating crawler identities by purpose — search, agent, training — stops being a best practice and becomes, in practice, a requirement for keeping access to content behind Cloudflare.