AI crawlers / CCBot
CCBot
Common Crawl · Common Crawl open corpus · Training crawler
The crawler for Common Crawl's open web archive. Its data is a widely used training corpus, so blocking it is a training choice, not a live-answer one.
AI crawlers / CCBot
Common Crawl · Common Crawl open corpus · Training crawler
The crawler for Common Crawl's open web archive. Its data is a widely used training corpus, so blocking it is a training choice, not a live-answer one.
At a glance
Your choice, this is a training control. Blocking it does not remove you from live AI answers.
CCBot is operated by Common Crawl, a nonprofit that maintains an open repository of web crawl data for research and development. Its user-agent is "CCBot/2.0 (https://commoncrawl.org/faq/)."
Common Crawl's archive is one of the most widely used sources of AI training data, so allowing or blocking CCBot is effectively a decision about your content appearing in training corpora. It does not decide live AI-answer eligibility.
Common Crawl warns that crawlers sometimes falsely identify as CCBot and recommends verifying with reverse DNS against its published IP ranges.
Add the matching group to the robots.txt at your site root.
Allow CCBot:
User-agent: CCBot
Allow: /Block CCBot:
User-agent: CCBot
Disallow: /Check what your live site actually allows with the Robots.txt + AI-Bot Tester.
Verified on 2026-09-04 against Common Crawl, CCBot. Confidence: documented, stated in the vendor’s own documentation.
Related crawlers
The next question
Being crawlable is step one. The AI Visibility Checker shows whether ChatGPT, Google AI Overviews, Perplexity and Gemini name and cite your brand, and who they name instead.
Keep going