The complete AI crawler list (2026)
Every real AI crawler from OpenAI, Anthropic, Perplexity, Google and Meta, verified directly against each company's own documentation, what each one actually does, and a generator for your own robots.txt.
- Every major AI company runs at least 2 separate crawlers: one for training, one for live answers. Blocking one does not block the other.
- Some "user" crawlers fetch a page live because a person asked about it, by the companies' own documentation, these can bypass robots.txt entirely.
- Build your own robots.txt below, pick what to allow and block, get a ready-to-paste file.
Why one company has 3 different bots
It's tempting to think "block the AI bot" is one decision. It isn't. Every major AI company splits its crawling into separate, independently-controllable bots because training a model and answering a live question are completely different jobs with different rules:
That third category matters more than people realize: by OpenAI's, Perplexity's, Anthropic's and Meta's own documentation, a live "user" fetcher exists specifically because someone typed a question that needs a specific page right now, and several of these are documented as ignoring robots.txt rules for that one request, the same way your own browser isn't blocked by someone else's robots.txt when you click a link.
Every crawler, by company
Check or uncheck any crawler below. It updates the generator further down the page.
OpenAI
Source: OpenAI's own crawler docs
Anthropic
Source: Anthropic's own crawler docs
Perplexity
Source: Perplexity's own crawler docs
Source: Google's own crawler docs
Amazon
Source: Amazon's own crawler docs
Meta
Source: Meta's own crawler docs
Your robots.txt
What "blocking" actually changes
- Block a training bot (GPTBot, ClaudeBot, Amazonbot, Meta-ExternalAgent) and your content stops feeding future model training. It has no effect on whether that assistant can still cite your page live today.
- Block a search/answer bot (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Meta-WebIndexer) and you directly remove yourself from that engine's live answers, this is the one that actually affects your AI visibility.
- Block a user fetcher (ChatGPT-User, Perplexity-User, Claude-User, Amzn-User, Meta-ExternalFetcher) and, per each company's own documentation, it may not even matter, several of these are built to fetch a specific page on a real person's behalf and can bypass robots.txt for that one request.
FAQ
Does blocking GPTBot remove me from ChatGPT's answers?
No. GPTBot is only for training OpenAI's models. Blocking it stops your content from being used in future model training, but ChatGPT can still cite your page live via OAI-SearchBot, which is a separate crawler with a separate robots.txt rule.
Which crawlers ignore robots.txt entirely?
Live, user-initiated fetchers, Perplexity-User, Meta-ExternalFetcher, and OpenAI's own docs note similar limits for ChatGPT-User, are built to fetch a specific page a person asked about, and by each company's own documentation may not fully honor robots.txt the way a normal crawl does.
Does Google-Extended affect my Google Search ranking?
No, by Google's own documentation. Google-Extended only controls whether your content can be used for Gemini training and grounding, Googlebot, the crawler that actually powers Search rankings, is entirely separate.
Should I just block every AI crawler?
That depends on what you're optimizing for. Blocking training crawlers keeps your content out of future model training but doesn't hurt live visibility. Blocking the search/answer crawlers does directly remove you from those engines' live answers.