TL;DR — Key Takeaways

  • Cloudflare’s new Disallow AI Training setting lets website owners restrict AI training while still allowing traditional search crawling.
  • The company’s new Accountable designation covers mixed-use crawlers that provide or commit to training opt-outs, AI-summary controls and greater visibility into content use.
  • Google and Apple already support separate training preferences through robots.txt, while Microsoft is targeting early 2027 for comparable domain-level support in Bing.

Cloudflare has rolled out new controls that let website owners restrict AI training without also blocking the crawlers that keep their content available to traditional search. The new “Disallow AI Training” setting went live today as part of an overhaul of Cloudflare’s AI crawler controls. It allows sites to keep search crawling enabled while publishing a no-training preference through robots.txt. Cloudflare said mixed-use crawlers that qualify under its new ‘Accountable’ criteria can continue indexing those sites for search while other training crawlers are blocked.

Mixed-use crawlers have been difficult for publishers to manage because the same bot may collect content for both a search index and AI model training. Blocking the crawler entirely could therefore protect content from training at the cost of search visibility. Cloudflare said fewer than 1% of sites on its network block search crawlers, while 17% use some mechanism to restrict AI training.

To qualify as Accountable, operators must offer or commit to a way for site owners to opt out of AI training through robots.txt or a comparable standard, as well as a separate way to opt out of AI-generated search summaries. They must also provide URL-level visibility into content use and publicly confirm that opting out of training will not affect a site’s ranking in traditional search, Cloudflare said.

Cloudflare has designated Applebot, Googlebot and Bingbot as Accountable, along with relevant crawlers from Amazon, Anthropic, Meta and OpenAI. Cloudflare said the latter four already separate their search and training crawlers, allowing it to block training crawlers without affecting search.

For Apple, Google and Microsoft, the implementation varies. Apple and Google already support robots.txt mechanisms that separate training preferences from search crawling. Microsoft currently lets publishers signal training preferences through Bing’s NOARCHIVE meta tag, but Cloudflare said Bing does not yet automatically honor the new domain-level no-training preference. Microsoft is currently building that support, targeting early 2027.

Cloudflare is also retiring its ‘Block AI Bots’ setting in favor of the separate Search, Training and Agent controls it introduced in July. Today’s update adds the new ‘Disallow AI Training’ option under the Training control. Cloudflare is also replacing Managed Robots.txt with Bot Preference Sync, which publishes applicable no-training preferences in robots.txt and applies a site owner’s choices across supported crawlers.

The company is also changing the recommended presets shown when customers add new domains. For sites that make money from advertising, it recommends allowing search crawlers, disallowing AI training and blocking agents on pages that serve ads. For sites that do not monetize through ads, Cloudflare recommends allowing search, AI training and agents. Customers can change any of those settings during onboarding or later.

Cloudflare plans to focus next on AI-generated search summaries. Operators designated as Accountable must provide or commit to a way for site owners to opt out of AI summaries, although those controls are currently set separately with each provider. The company said that by early next year, it plans to let site owners manage those preferences in one place and specify how much of their content can appear in summaries. Read more about the changes in this Cloudflare blog post.