Robots.txt AI Crawler Checker
See which AI and search crawlers your robots.txt allows or blocks — GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Googlebot, Google-Extended, Bingbot and Applebot — and whether each one is named explicitly or just falls under the default rule.
What the checker shows
We fetch your live robots.txt and evaluate it the way major crawlers do (RFC 9309). For each crawler you get one of three answers — and we never blur them:
- Explicitly allowed — robots.txt has a group for this crawler, and it may fetch your homepage.
- Explicitly blocked — robots.txt has a group for this crawler, and it disallows your homepage.
- No specific rule found — nothing names this crawler, so it follows the
*group (or, if there isn't one, everything is allowed).
We also show whether some paths are disallowed even when the homepage is allowed, and list any sitemaps your robots.txt declares.
The crawlers that matter for AI search
- OAI-SearchBot — used to show websites in ChatGPT search results. Blocking it can keep your pages out of ChatGPT's cited sources.
- GPTBot — OpenAI's crawler for content that may be used to improve its models. It is separate from OAI-SearchBot, so you can allow one and block the other.
- ClaudeBot — Anthropic's crawler.
- PerplexityBot — used by Perplexity to find and link websites in its answers.
- Googlebot — Google Search. Blocking it removes you from Google.
- Google-Extended — not a crawler but a control token: it decides whether Google may use your content for Gemini models and grounding. It does not affect Google Search rankings or crawling.
- Bingbot and Applebot — Bing (which also powers other services, including Copilot) and Apple (Siri and Spotlight).
How robots.txt rules are read
A crawler looks for a group that names its user-agent token. If one exists, it uses only that group and ignores *. If none exists, it uses *. Within a group, the longest matching rule wins, and Allow wins a tie.
Example — blocks GPTBot, allows everyone else
- User-agent: GPTBot
- Disallow: /
- User-agent: *
- Allow: /
Here GPTBot is explicitly blocked; OAI-SearchBot, ClaudeBot and the rest have no specific rule and follow *, so they're allowed.
robots.txt is a request, not access control. Well-behaved crawlers follow it; it doesn't hide pages from people or from crawlers that ignore it.
Should you block AI crawlers?
It's a business decision. If you want AI search to recommend and cite you, the search-focused crawlers (OAI-SearchBot, PerplexityBot) need access to the pages you want cited. Some sites block training-focused crawlers while allowing search crawlers — robots.txt lets you do both.
Access is only the first step. To see whether AI actually mentions you, run the free AI visibility checker, or check other technical signals with the AI readiness checker.
Frequently asked questions
Does blocking Google-Extended remove me from Google Search?
+
No. Google-Extended only controls use of your content for Gemini models and grounding. Google Search crawling is controlled by Googlebot.
My robots.txt doesn't mention GPTBot. Is it allowed?
+
It follows your * group. If * allows your pages, GPTBot is allowed; if there's no * group either, everything is allowed. The checker shows this as "No specific rule found".
What if I have no robots.txt?
+
A missing robots.txt (404) means crawlers treat everything as allowed. A robots.txt that returns a server error is different: many crawlers pause or treat the site as blocked until it works again.
Do you store the sites I check?
+
No. We only keep a salted hash of your IP address and a timestamp for rate limiting — not the website you entered.
Start with a free check, then track it over time
The free check asks ChatGPT 3 real customer questions about your business. Pro monitors up to 10 questions on ChatGPT and Perplexity every 7 days — Pro is launching soon.