Free tool

AI crawler access check

ChatGPT, Claude, Perplexity and Gemini each send a named crawler. One line in robots.txt decides whether they can read your site. Paste your address and see which ones are turned away.

The domain only. No login, nothing is stored.

Want this running 24/7 on your account?

Connect Meta in 60 seconds. First scan is free.

What does this check actually do?

It makes one request for your robots.txt and reads it the way a crawler does. For each of twenty named AI agents it works out whether the file allows or refuses a request for your home page, and it shows you the exact line that decided it.

Nothing else is fetched. Your pages are not crawled, nothing is stored, and no account is involved. The whole answer comes from one public text file that anybody can read.

How robots.txt rules are resolved

Crawlers do not simply read the first rule they find. A group naming an agent by name replaces the wildcard group entirely, and inside a group the longest matching path pattern wins, with a tie going to Allow. That is why a file that starts Disallow: / under User-agent: * can still be perfectly open to GPTBot further down.

Getting that wrong in either direction is expensive, so this check follows the rule resolution in RFC 9309 rather than pattern-matching the file. It reports three states, never two: allowed, blocked, and not known. An unreachable robots.txt gets the third one.

What to do with the result

If an answer engine is blocked and you want to be quoted, add the lines the tool prints and redeploy. If everything is allowed, the next question is whether an assistant can make sense of what it finds, which is what the llms.txt generator is for.

Keep reading

Last updated August 31, 2026. All 17 calculators and AI helpers are listed on the free tools hub.

FAQ

Common questions

Which crawlers does this check?

Twenty named agents, in three groups. Answer engines that fetch a page to answer a question right now: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Bingbot and YouBot. Training crawlers that build the corpus a model learns from: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Amazonbot, Bytespider, cohere-ai, Diffbot, meta-externalagent and Timpibot. And FacebookBot, which renders link previews.

What is the difference between blocking an answer engine and blocking a training crawler?

Blocking an answer engine costs you today. When somebody asks ChatGPT about your category, the engine fetches candidate pages and yours is refused, so it cites a competitor instead. Blocking a training crawler costs you later: your pages are absent from the corpus the next model generation learns from, and the effect shows up whenever that model ships.

My robots.txt did not respond. Does that mean I am blocked?

No, and this tool will not say so. A file that times out or returns a server error tells us nothing, so every crawler is reported as not known rather than blocked. A missing robots.txt is different again: a 404 means no rule exists, which means everything is allowed.

Should I allow every AI crawler?

That is a business decision, not a technical one. If you want to be quoted in AI answers, the answer engines have to be able to fetch the page. Training crawlers are a longer bet and some publishers deliberately refuse them. The point of this check is that the choice should be one you made, rather than one a default robots.txt made for you.