GapCite
GapCite Blog ยท AI visibility basics

The complete AI crawler list (2026)

Every real AI crawler from OpenAI, Anthropic, Perplexity, Google and Meta, verified directly against each company's own documentation, what each one actually does, and a generator for your own robots.txt.

By Joe WangUpdated September 202620 crawlers, verified at the source
The short version
  1. Every major AI company runs at least 2 separate crawlers: one for training, one for live answers. Blocking one does not block the other.
  2. Some "user" crawlers fetch a page live because a person asked about it, by the companies' own documentation, these can bypass robots.txt entirely.
  3. Build your own robots.txt below, pick what to allow and block, get a ready-to-paste file.

Why one company has 3 different bots

It's tempting to think "block the AI bot" is one decision. It isn't. Every major AI company splits its crawling into separate, independently-controllable bots because training a model and answering a live question are completely different jobs with different rules:

Trains future models Powers live search/answers Fetches a page because a user asked

That third category matters more than people realize: by OpenAI's, Perplexity's, Anthropic's and Meta's own documentation, a live "user" fetcher exists specifically because someone typed a question that needs a specific page right now, and several of these are documented as ignoring robots.txt rules for that one request, the same way your own browser isn't blocked by someone else's robots.txt when you click a link.

Every crawler, by company

Check or uncheck any crawler below. It updates the generator further down the page.

ToolRobots.txt generator
Checked = allowed. Unchecked = blocked. Defaults reflect what most businesses actually want.

OpenAI

Source: OpenAI's own crawler docs

Anthropic

Source: Anthropic's own crawler docs

Perplexity

Source: Perplexity's own crawler docs

Google

Source: Google's own crawler docs

Amazon

Source: Amazon's own crawler docs

Your robots.txt


      

What "blocking" actually changes

FAQ

Does blocking GPTBot remove me from ChatGPT's answers?

No. GPTBot is only for training OpenAI's models. Blocking it stops your content from being used in future model training, but ChatGPT can still cite your page live via OAI-SearchBot, which is a separate crawler with a separate robots.txt rule.

Which crawlers ignore robots.txt entirely?

Live, user-initiated fetchers, Perplexity-User, Meta-ExternalFetcher, and OpenAI's own docs note similar limits for ChatGPT-User, are built to fetch a specific page a person asked about, and by each company's own documentation may not fully honor robots.txt the way a normal crawl does.

Does Google-Extended affect my Google Search ranking?

No, by Google's own documentation. Google-Extended only controls whether your content can be used for Gemini training and grounding, Googlebot, the crawler that actually powers Search rankings, is entirely separate.

Should I just block every AI crawler?

That depends on what you're optimizing for. Blocking training crawlers keeps your content out of future model training but doesn't hurt live visibility. Blocking the search/answer crawlers does directly remove you from those engines' live answers.

Sources