Skip to content

Reference · AI crawlers

The AI crawler reference

Every major AI crawler, what it actually controls, and how to allow or block it. Each fact is checked against the vendor’s own documentation and carries a last-verified date. The one rule most guides get wrong: the crawlers that get you cited are not the same as the ones that train a model.

How to read this

An AI engine will only cite a page it can reach. But “allow the AI crawlers” is too blunt, the crawlers split into three jobs, and confusing them is the most common way sites accidentally opt out of citations or waste effort blocking the wrong bot.

  • Discovery / citation crawlers decide whether you can appear in a live AI answer or AI search result. Allow these to be eligible.
  • User-initiated fetchers visit a page only when a person asks the assistant about it. They usually ignore robots.txt because they act on a direct request.
  • Training crawlers collect content to train or ground a model. Blocking one opts you out of training, it does not remove you from live AI answers.

Frequently asked questions

Which AI crawlers decide whether I get cited?

The discovery / citation crawlers: OAI-SearchBot (ChatGPT), Googlebot (Google Search and AI Overviews), PerplexityBot (Perplexity), Claude-SearchBot (Claude), and Bingbot (Bing and Copilot). Blocking any of these makes you ineligible to appear in that engine’s live answers.

Does blocking GPTBot or Google-Extended remove me from AI answers?

No. GPTBot and Google-Extended are training controls. Blocking them opts your content out of model training; it does not remove you from live AI answers. Google states plainly that Google-Extended does not affect inclusion in Google Search.

What is a user-initiated fetcher?

An agent like ChatGPT-User, Perplexity-User or Claude-User that visits a page only when a person asks the assistant about it. These often disregard robots.txt because they act on a direct user request rather than crawling in the background.

How current is this data?

Every entry carries a "last verified" date and a link to the vendor’s own documentation. This registry was last verified on 2026-09-04. The machine-readable version is at /data/ai-crawlers.json.

The next question

Your page is AI-ready. But does AI actually cite you?

Being crawlable is step one. The AI Visibility Checker shows whether ChatGPT, Google AI Overviews, Perplexity and Gemini name and cite your brand, and who they name instead.

Join the waitlist

Keep going

More free tools

All tools →