Free tool by TK WebHosts

AI Crawler Access Checker

Can ChatGPT, Claude and Perplexity actually read your website? Check which search and AI crawlers your robots.txt allows — separated into training bots and the live agents that decide whether AI answers cite you.

Why it matters

Training and citation are separate decisions.

Training crawlers

GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider and Meta collect content to train future models. Blocking them is a content-licensing stance.

Answer-time crawlers

OAI-SearchBot, ChatGPT-User, Claude-SearchBot and PerplexityBot fetch your page when a user asks a question — they are how AI assistants cite and recommend you.

Search crawlers

Googlebot and Bingbot remain the foundation. Blocking them by accident is a site-killer — this checker flags it in red immediately.

AI crawler FAQs

robots.txt, training bots and answer-time agents.

What is the difference between GPTBot and ChatGPT-User?

GPTBot collects pages to train OpenAI's future models. ChatGPT-User and OAI-SearchBot are the live agents: they fetch and cite your pages when someone asks ChatGPT a question right now. Blocking GPTBot only opts out of training — it does not remove you from ChatGPT's live answers. The same split applies to Anthropic (ClaudeBot vs Claude-User/Claude-SearchBot).

If I block AI training crawlers, do I disappear from AI search answers?

No. Training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended) and answer-time crawlers (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot) are separate user-agents with separate robots.txt rules. Many publishers block training while staying open to citation — this checker shows both groups separately so you can see exactly which policy you have.

Should I block AI crawlers?

It depends on your goal. If you sell content itself, blocking training crawlers is a reasonable policy. If you sell services or products, being cited in AI answers is free distribution — blocking the answer-time crawlers means competitors get recommended instead of you.

What happens if my site has no robots.txt?

Every crawler is allowed by default. A missing robots.txt is not an error — but a robots.txt that returns a server error (500/503) is serious: Google treats it as “do not crawl the site at all”.

Does robots.txt actually stop crawlers?

It is a convention, not a lock. The major operators listed here — Google, Microsoft, OpenAI, Anthropic, Perplexity, Apple, Meta — publicly commit to honouring it. It will not stop a scraper that ignores the rules.

Want the whole picture?

Run the full Website SEO Checker — 10 scored categories, a prioritised to-do list, and a shareable report.

Run a Free SEO Check