Check in one click whether GPTBot, ClaudeBot, Google-Extended and other AI crawlers can read your pages — or whether your robots.txt is quietly blocking them.
AI crawlers are bots run by companies like OpenAI, Anthropic and Google to read public web pages and feed their AI models and answer engines. If they can read your site, your brand can show up in ChatGPT, Claude, Perplexity and AI Overviews.
Your robots.txt file — a plain text file at the root of your domain — tells these bots what they may and may not access. A single Disallow line can quietly remove you from AI answers.
Type your website address. We fetch your public /robots.txt file.
We match each major AI crawler against your User-agent and Disallow rules.
Get a clear allowed / blocked status for every AI crawler in seconds.
Your robots.txt file is a plain-text file at the root of your domain that tells automated bots which parts of your site they may and may not fetch. For two decades it governed search engine spiders like Googlebot. Today it has a second, equally important job: deciding which AI crawlers — the bots that feed ChatGPT, Perplexity, Claude and Google’s generative answers — are allowed to read your content.
AI vendors run distinct user-agents for distinct purposes. OpenAI uses GPTBot to gather training data, OAI-SearchBot to power live search answers in ChatGPT, and ChatGPT-User for on-demand fetches when a user asks ChatGPT to open a link. Anthropic runs ClaudeBot and the older anthropic-ai. Perplexity uses PerplexityBot. Google uses Google-Extended as an opt-out token for Gemini training, separate from its Googlebot search crawler. Common Crawl’s CCBot feeds many downstream models, and Bytespider (ByteDance), Amazonbot and Applebot-Extended round out the list.
The key insight is that these bots do different things, so a blanket block is rarely the right move. Blocking a training crawler protects your content from being absorbed into a model, but blocking a live-retrieval bot like OAI-SearchBot or PerplexityBot can quietly remove you from the AI answers your buyers see — the AI equivalent of deindexing yourself from Google. Most brands want the opposite: maximum visibility in AI answers, with selective control over bulk scraping.
robots.txt is a directive, not a hard wall. Reputable crawlers from OpenAI, Anthropic, Perplexity and Google honor it, but it relies on the bot choosing to obey. It is the standard, low-friction way to express your preferences, and pairing it with an llms.txt file gives well-behaved AI systems a clean map of the content you most want them to use.
A practical sequence for auditing what you allow today, opening the doors that grow AI visibility, and closing the ones you do not want — without accidentally hiding from ChatGPT and Perplexity.
Run your domain through this checker to see which AI user-agents you allow, block or partially block today. Many sites unknowingly block GPTBot or PerplexityBot because of a copied-and-pasted rule or a security plugin default.
Separate live-retrieval bots (OAI-SearchBot, ChatGPT-User, PerplexityBot) from training bots (GPTBot, Google-Extended, CCBot, ClaudeBot). Most brands allow retrieval bots for visibility and make a deliberate choice on training bots.
To appear in ChatGPT and Perplexity answers, explicitly allow OAI-SearchBot, ChatGPT-User and PerplexityBot. If you want your content used in answers and training, allow GPTBot and ClaudeBot too.
If bandwidth, bulk scraping or unattributed reuse is a concern, add Disallow rules for aggressive crawlers such as Bytespider, CCBot or Amazonbot while leaving the answer-engine bots allowed.
Use one User-agent block per bot, name the user-agent verbatim, and use Allow: / or Disallow: / under it. A trailing rule that is too broad can block far more than you intend, so keep each directive scoped.
Re-run this checker after editing robots.txt to confirm each bot reads as you intended. Verify the file returns HTTP 200 at /robots.txt and is not blocked by a firewall, CDN or WAF rule.
Publish an llms.txt that lists your most important pages and documentation. It complements robots.txt by pointing well-behaved AI systems at the content you most want cited.
AI vendors add and rename user-agents regularly. Re-check quarterly, watch your server logs for new AI bots, and confirm a CDN update or new plugin has not silently changed your rules.
The vocabulary of controlling AI bots, in plain English.
If buyers research you through AI assistants, your robots.txt now decides whether they can read you.
Decide deliberately whether to allow training crawlers like GPTBot and CCBot while keeping answer-engine bots such as OAI-SearchBot and PerplexityBot allowed so your reporting still gets cited.
Confirm ChatGPT and Perplexity can read your docs and feature pages, so when buyers ask AI for the best tool in your category, your site is available to be quoted.
Make sure product and category pages are crawlable by AI shopping answers, and selectively limit aggressive scrapers that strain your servers without sending traffic.
Audit a client’s robots.txt in seconds, spot accidental blocks of GPTBot or PerplexityBot, and turn the fix into a clear, measurable AI-visibility win.
Access is step one. See whether AI engines actually recommend your brand — and how to make them.
Get My Free RAIVE Score