Training crawlers
GPTBot, ClaudeBot, Google-Extended, CCBot and others collect content to train future models. Blocking them is a content-licensing stance.
Free tool by TK WebHosts
Can ChatGPT, Claude and Perplexity actually read your website? Check which search and AI crawlers your robots.txt allows — separated into training bots and the live agents that decide whether AI answers cite you.
| Crawler | Operator | Access |
|---|
Add this to your robots.txt to open the answer-time crawlers (you can keep training bots blocked):
GPTBot, ClaudeBot, Google-Extended, CCBot and others collect content to train future models. Blocking them is a content-licensing stance.
OAI-SearchBot, ChatGPT-User, Claude-SearchBot and PerplexityBot fetch your page when a user asks a question — they are how AI assistants cite you.
Googlebot and Bingbot remain the foundation. Blocking them by accident is a site-killer — this checker flags it in red immediately.
GPTBot collects pages to train future models. ChatGPT-User and OAI-SearchBot fetch and cite your pages when someone asks ChatGPT a question right now. Blocking GPTBot only opts out of training.
No. Training crawlers and answer-time crawlers are separate user-agents with separate robots.txt rules. Many publishers block training while staying open to citation.
Every crawler is allowed by default. A missing robots.txt is not an error — but one that returns a server error is serious: Google treats it as "do not crawl".
It is a convention, not a lock. The major operators listed here publicly commit to honouring it; it will not stop a scraper that ignores the rules.
Run the full SEO Checker — ten scored categories, a prioritised to-do list, and a shareable report.
Run a Free SEO Check