TL;DR — Key Takeaways

  • AI crawler identities have split by function, with separate bots for training, search indexing and live page retrieval.
  • robots.txt rules can produce unintended results, especially when wildcard groups block newer AI crawlers without anyone explicitly deciding to do so.
  • Crawler policy now spans robots.txt, content-use signals and infrastructure layers such as CDNs or hosting platforms, so teams need to test actual access from outside.

AI crawlers have split by job. One agent trains a model, another builds the index behind an assistant’s search, and a third fetches a page live. A single rule aimed at the AI bots hits some and misses others.

In a published sample of 3,000 domains, 196 block OpenAI’s training crawler while allowing its search crawler. None do the reverse.

Around half the sites closed to a citation crawler never wrote an AI rule at all. A wildcard group written for something else is doing it.

Cloudflare’s Content Signals Policy adds a second decision to the same file: not who may fetch, but what the content may be used for afterwards.

Crawler access is now decided in three places. Only one of them is in your repository.

For twenty years, robots.txt answered one question: May this crawler fetch this path? That question is now the smaller half of what the file decides.

The larger half is what happens to the content after it is fetched, and it is where most teams have made no decision at all, which is itself a decision.

The Crawlers Split, and the Names Stopped Matching the Jobs

OpenAI documents three separate agents. GPTBot collects training data. OAI-SearchBot builds the index behind ChatGPT search. ChatGPT-User fetches a page live when a user’s question triggers browsing.

Anthropic documents the same separation, and states plainly that blocking ClaudeBot does not block Claude-SearchBot or Claude-User. Google states in writing that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”.

Three vendors, three sets of tokens, one shared implication: A single rule aimed at the AI bots will hit some jobs and miss others, and the one it usually hits is not the one the team was worried about.

What 3,000 Domains Actually Do

We sampled 3,000 domains from the Tranco top million and parsed their robots.txt files, publishing the sampling frame and the code. 1,453 returned a usable file, a coverage rate of 48.4%, which is reported rather than hidden.

In that sample, 196 domains block OpenAI’s training crawler while allowing its search crawler. Zero does the reverse. As a stated position that is coherent: do not train on me, but do cite me. What makes it interesting is the other side of the ledger.

12.04% of the sample is closed to at least one citation crawler, and about half of those never wrote a rule about AI at all. They were caught by a wildcard block written for something else entirely. Against that, 73.41% of GPTBot blocks named GPTBot explicitly. Training blocks are decisions. Citation blocks are mostly accidents.

Two smaller patterns from the same sample are worth a minute of anyone’s time.

  • 17.00% block Google-Extended while leaving Googlebot open. By Google’s own wording, that does not remove those pages from AI Overviews, because those serve from the Search index. Teams that set this rule to stay out of AI answers did not achieve what they intended.
  • 3.03% disallow everything for the wildcard agent and then allow Googlebot by name. That pattern was reasonable when Googlebot was the only crawler that mattered. It now silently excludes every engine that appeared afterwards.

Content-Signal is the Second Decision

Cloudflare’s Content Signals Policy, now applied across millions of domains on its network, adds a machine-readable line to robots.txt expressing preferences about use rather than access. It carries three signals: Search, meaning use for a traditional search index with links and short excerpts; ai-input, meaning real-time use in a generated answer; and ai-train, meaning training or fine-tuning. Each is yes, no, or absent, and absent means no preference expressed.

Two honest caveats belong with it. It is a preference expressed to parties who may or may not honour it, and it is not a standard. Cloudflare says so itself.

That does not make it pointless. It makes it a stated position, in public, in a file anyone can read, which is a different thing from an access rule and is worth setting deliberately rather than inheriting. A site that allows a crawler while expressing no preference about use has answered the first question and left the second blank.

Parse it by the Spec, Not by Grep

Any audit of this at scale runs into the same trap. Searching a robots.txt for a bot’s name gets the answer wrong in both directions. A file can name GPTBot inside a group that allows it. A file can block every crawler through a wildcard group while naming no AI bot at all. Our parser implements the group matching rules in RFC 9309 for exactly that reason, and the second case is where the accidental blocks above come from.

The same applies to the one-line audits that circulate internally. “We do not block any AI crawlers” is usually derived from a search for names, while the wildcard group three lines above is what actually governs.

The Layer Below the File

There is a failure mode that no robots.txt audit can see, because the block sits underneath it.

Earlier this month we measured a site on a large shared host receiving HTTP 429 for GPTBot on cache misses, while seven other crawler identities from the same IP in the same minutes received 200. The host confirmed in writing that the rate limit is intentional, applied at the infrastructure level, and not configurable per domain. The site’s own robots.txt welcomed the crawler. The request never reached the site, and the same test can be run against any host in about ten minutes.

The general point is not about one host. It is that crawler access is now decided in at least three places: the file, the edge or CDN in front of the origin, and the platform the site runs on. Only the first is version controlled, and only the first is what anybody checks.

Four Checks Worth Running This Quarter

  • Read your robots.txt as a parser would, by group, not by name search. Confirm which groups a given agent actually matches.
  • List the agents by job, not by vendor. For each vendor, know which token trains, which one indexes for the assistant’s search, and which one fetches live. Then decide each one separately.
  • Set Content-Signal deliberately, including choosing to leave it unset, and write down why. Inheriting a default from a platform update is not a position.
  • Test from outside. Request a cache-missing page with each agent’s user agent string and record the status code. If the answer differs from what your file says, the block is in the edge or the host, not in your repository.

None of this is expensive. It is an afternoon, once, and then a line in the change log the next time someone edits the file.

The reason to do it now is narrow and practical. The teams in that sample who are invisible to the engines that answer questions mostly did not choose to be. They wrote one rule, years ago, for a different web.

Frequently Asked Questions

Why is one AI crawler rule no longer enough?
Different crawler identities now perform different jobs. Blocking a training bot may not block a search-indexing or live-browsing bot from the same vendor.
What is Cloudflare Content Signals?
It adds machine-readable preferences describing whether content may be used for traditional search, AI-generated answers or AI training. It expresses a preference rather than technically preventing access.
Why can robots.txt audits be misleading?
A crawler may be affected by wildcard rules even if its name never appears in the file. It may also be blocked at the CDN or hosting layer despite robots.txt explicitly allowing it.

TECHSTRONG AI PODCAST

SHARE THIS STORY