Blog · Sep 22, 2026
robots.txt and AI crawlers
The problem
Many sites block AI crawlers without knowing it. Sometimes it's a single Disallow: / under a wildcard rule, left over from a staging deploy. Sometimes a security plugin blocks "bots" aggressively. Either way, the result is the same: ChatGPT, Gemini, and Perplexity can't read your site, so they don't mention it.
Know the crawlers
The bots that matter in 2026:
- GPTBot — OpenAI's training crawler
- ChatGPT-User / OAI-SearchBot — OpenAI's live answer crawlers
- ClaudeBot / anthropic-ai — Anthropic's crawlers
- PerplexityBot — Perplexity's crawler
- CCBot — Common Crawl (feeds many AI training sets)
- Google-Extended — controls use of your content in Google's AI products
- Applebot-Extended — Apple's AI training control
- Bytespider — ByteDance's crawler
Audit yours in 30 seconds
Open yoursite.com/robots.txt in a browser. Look for User-agent lines naming any bot above, followed by Disallow: /. Also check the wildcard block — User-agent: * with Disallow: / blocks everything, including AI crawlers.
The business decision
Blocking isn't always wrong — it's a tradeoff:
- Allow if you want AI systems to cite and recommend you. For most businesses, visibility wins.
- Block if your content is the product (paywalled journalism, proprietary data) and you don't want it in training sets.
Note the nuance: some bots separate training from live answers. Blocking GPTBot stops training use; blocking OAI-SearchBot stops ChatGPT from citing you in answers. You can allow one and block the other.
Example configs
Allow the crawlers that drive visibility:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
Block AI training while allowing live answers:
User-agent: GPTBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
The wildcard trap
This is the #1 cause of accidental blocking we see in scans:
User-agent: *
Disallow: /
Meant for staging, copied to production, forgotten. Every bot — Google, Bing, and all AI crawlers — obeys it. If your site is live and this is in your robots.txt, fix it today.
How CrawlReady checks it
Every scan fetches your robots.txt and tests each major AI crawler against it individually — not just the wildcard. The report names exactly which bots are blocked and shows the rule doing it, so the fix takes minutes.