AI CRAWLER CHECK

Who is actually training
on your content?

Enter your domain — the check reads your robots.txt and shows which AI crawlers you explicitly allow or block, and which ones you simply let through by default. In two seconds.

Background

Who trains on your website? What the check inspects

The check reads your robots.txt and server responses and shows which AI crawlers may access your content — GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Gemini), CCBot (Common Crawl), PerplexityBot and more. For each bot you see: allowed, blocked or unregulated.

This is a strategic decision, not just a technical one. Block everything and your content appears less in AI answers — which are becoming the first touchpoint for more and more buying decisions. Allow everything and your content also feeds training data. Many businesses deliberately differentiate: search and AI answers yes, training no.

The result shows you the current state and where your robots.txt deviates from your intent. If you want to work out a policy that fits your business model, we can do that in a free first call.

Frequently asked questions

Does every bot respect robots.txt?

The major providers (OpenAI, Anthropic, Google) document their crawlers and follow the rules. Some smaller scrapers ignore robots.txt — only server-side measures help against those.

Should I block AI crawlers?

Depends on your goal. If you want to be found via AI search, allow answer crawlers. If you protect exclusive content, block training specifically. Both at once is possible — the bots are separated by purpose.

What's the difference between training and search?

Training crawlers (e.g. GPTBot) collect content to train models. Search crawlers (e.g. OAI-SearchBot) fetch content live to cite and link it in answers — the latter sends you visitors directly.