Controlling AI crawler access in robots.txt
By Nalyzed · Published 2026-10-07 · Updated 2026-10-07 · 6 min read
AI crawlers split into three jobs: collecting training data, fetching pages to answer a live question, and fetching a page a user explicitly pasted. Blocking the training crawlers has no effect on whether you are cited; blocking the search crawlers removes you from AI answers entirely. Most sites that are invisible in AI search got there by accident, with one over-broad robots.txt rule.
Three kinds of crawler, three different decisions
Treating every AI user-agent as one category is what produces the wrong outcome. The operators document distinct tokens for distinct purposes, and they behave differently.
- Training: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended, Applebot-Extended, CCBot (Common Crawl), Bytespider (ByteDance), meta-externalagent (Meta), cohere-ai, AI2Bot, Timpibot. These build datasets. Blocking them does not affect citations.
- Search and answer: OAI-SearchBot (OpenAI), Claude-SearchBot (Anthropic), PerplexityBot, Amazonbot, DuckAssistBot, YouBot, PetalBot. These fetch pages to construct an answer. Blocking them removes you from that answer.
- User-initiated: ChatGPT-User, Claude-User, MistralAI-User. These fetch a page because a person asked about that specific URL. Blocking them breaks the case where someone pastes your link.
The rules most sites get wrong
robots.txt syntax is unforgiving in ways that are easy to miss, and the failures are silent — nothing reports that your rule did not apply.
- A group runs from its User-agent lines to the next User-agent line. Blank lines and comments inside it do not end it, but some hand-written files assume they do.
- Group names match as a prefix of the crawler token, case-insensitively, and the longest match wins. A group for Googlebot also covers Googlebot-News.
- A crawler obeys only the single most specific group that matches it. Adding a specific group means the wildcard rules no longer apply to that crawler at all — including any Disallow lines you still wanted.
- An empty Disallow means allow everything. Disallow: / means block everything.
- robots.txt controls fetching, not indexing or snippets. To control excerpts use the robots meta tag or the X-Robots-Tag header.
A worked example: opt out of training, stay citable
This is the configuration most publishers actually want, and the one that is easiest to get subtly wrong. Note that each specific group has to repeat any shared rules, because a crawler reads only its own group.
- Add a group for each training crawler with Disallow: / — GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, meta-externalagent.
- Add a group for each search and user crawler with Disallow: left empty, plus whatever private paths you block for everyone — OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User.
- Keep your existing User-agent: * group unchanged for everything else.
- Make sure no robots meta tag or X-Robots-Tag on those pages sends nosnippet or max-snippet:0, which would forbid the excerpt an answer needs.
- Verify each token individually rather than assuming the file reads the way you intended.
What blocking does not achieve
robots.txt is a convention honoured by well-behaved crawlers. It is not an access control mechanism: it does not stop a scraper that ignores it, and it does not retroactively remove content from a model already trained on it.
It also does not stop a person from pasting your page into a chat. If content genuinely must not leave your site, it needs authentication, not a robots directive.
Frequently asked questions
- Will blocking AI crawlers hurt my Google rankings?
- Blocking Google-Extended does not affect Google Search ranking or indexing — it only opts you out of Gemini and Vertex AI training. Blocking Googlebot itself is what removes you from Search, and the two are separate tokens.
- Why did adding a rule for one crawler change its behaviour unexpectedly?
- Because a crawler obeys only the most specific matching group. The moment you create a group for it, your User-agent: * rules stop applying to it, so any Disallow lines you still wanted must be repeated inside the new group.
- Do AI crawlers respect robots.txt at all?
- The major documented ones do, and the operators publish their tokens and IP ranges so you can verify. Undocumented scrapers are a different problem that robots.txt was never able to solve.
- How do I stop AI engines from quoting my page without blocking them?
- Allow the crawler but send nosnippet, or max-snippet with a small value, via a robots meta tag or the X-Robots-Tag header. The page stays indexable while excerpts are restricted — though that usually also removes you from answers.
Check your own site
More guides
Published by Nalyzed. Scoring rules and limits are documented on the methodology page.