Skip to content

AI crawlers / CCBot

CCBot

Common Crawl · Common Crawl open corpus · Training crawler

The crawler for Common Crawl's open web archive. Its data is a widely used training corpus, so blocking it is a training choice, not a live-answer one.

At a glance

Affects live AI answersNo
Affects search inclusion,
Used for model trainingYes
robots.txt tokenCCBot
Source confidencedocumented · verified 2026-09-04

Your choice, this is a training control. Blocking it does not remove you from live AI answers.

What CCBot is

CCBot is operated by Common Crawl, a nonprofit that maintains an open repository of web crawl data for research and development. Its user-agent is "CCBot/2.0 (https://commoncrawl.org/faq/)."

Common Crawl's archive is one of the most widely used sources of AI training data, so allowing or blocking CCBot is effectively a decision about your content appearing in training corpora. It does not decide live AI-answer eligibility.

Common Crawl warns that crawlers sometimes falsely identify as CCBot and recommends verifying with reverse DNS against its published IP ranges.

Controlling CCBot in robots.txt

Add the matching group to the robots.txt at your site root.

Allow CCBot:

User-agent: CCBot
Allow: /

Block CCBot:

User-agent: CCBot
Disallow: /

Check what your live site actually allows with the Robots.txt + AI-Bot Tester.

Source

Verified on 2026-09-04 against Common Crawl, CCBot. Confidence: documented, stated in the vendor’s own documentation.

The next question

Your page is AI-ready. But does AI actually cite you?

Being crawlable is step one. The AI Visibility Checker shows whether ChatGPT, Google AI Overviews, Perplexity and Gemini name and cite your brand, and who they name instead.

Join the waitlist

Keep going

More free tools

All tools →