Skip to content

Crawl control

Free robots.txt tester

Fetch any site's robots.txt or paste your own, then test a URL path against Googlebot, Bingbot, GPTBot or any crawler you name. You get an allowed or blocked verdict and the rule that decided it, highlighted in the file.

  • Free
  • No signup
  • Test before you deploy

How it works

Fetch the file, pick a crawler, test a path

01

Fetch or paste the file

Enter a domain and Rank.ai fetches its robots.txt for you, or paste a draft you haven't deployed yet.

02

Choose a crawler and a path

Pick Googlebot, Googlebot-Image, Bingbot, GPTBot, ClaudeBot or PerplexityBot, or type any user agent, then enter the path you want to test.

03

See the verdict and the rule

The tester applies the file the way RFC 9309 describes and shows allowed or blocked, with the matching rule highlighted by line number.

Why it matters

How robots.txt rules get applied

Search Console’s robots.txt report shows the robots.txt files Google found for a site you have verified, when it last crawled them and any errors. It doesn’t test a path against a draft you haven’t deployed, or a site you can’t verify. This tool does both, applying the matching rules defined in RFC 9309.

The precedence rules often surprise people. A crawler follows exactly one group: the one whose user agent token best matches its own name, falling back to User-agent: * only when nothing else matches. Rules in other groups don’t apply to it. Within that group, the rule with the longest path pattern wins, whatever order the rules appear in, and when an Allow and a Disallow of equal length both match, Allow wins. That is why Allow: /blog/ beats Disallow: / for blog URLs.

Some mistakes are worth testing for directly. A Disallow: / left over from staging stops the whole site being crawled. Paths are case sensitive, so Disallow: /Admin/ does nothing for /admin/. Wildcards match widely, so Disallow: /*? blocks every URL with a query string, campaign links included. And robots.txt controls crawling only; a blocked page can still appear in search results if other sites link to it.

AI companies crawl with their own bots, and most separate search from training: OpenAI’s OAI-SearchBot and PerplexityBot put pages in AI search answers, while GPTBot and ClaudeBot collect training data. Rules written years ago can block them without anyone meaning to. The AI crawler checker shows how your file treats each of them. While you are auditing crawlability, check your structured data with the schema validator and your titles with the title tag checker.

Frequently asked

How is this different from Search Console's robots.txt report?
Google's robots.txt report shows the robots.txt files Google found for a verified site, when it last crawled them and any warnings or errors, and lets you request a recrawl. It doesn't test a path against a draft file or a site you can't verify, which is what this tool does.
Which rule wins when both an Allow and a Disallow match?
The rule with the longest path pattern wins, counted by characters in the pattern. If an Allow and a Disallow of exactly equal length both match, Allow wins the tie, and the order in the file doesn't matter. With Disallow: /shop/ and Allow: /shop/sale/, a URL under /shop/sale/ is allowed.
Does Disallow remove a page from Google's index?
No. Disallow stops crawling, and that is all it does. A page Google can't crawl can still be indexed and shown if other pages link to it, usually with no snippet. To keep a page out of the index, allow crawling and add a noindex tag or header.
Are robots.txt paths case sensitive?
Yes. Paths match exactly, so Disallow: /Private/ doesn't block /private/. User agent tokens are the opposite: they match regardless of case, so user-agent: googlebot and User-agent: Googlebot are the same.
Where does robots.txt have to live?
At the root of the exact host, such as https://www.example.com/robots.txt. It applies per protocol, host and port: the file on www.example.com doesn't cover shop.example.com or the http version, and a robots.txt in a subfolder is ignored.

See if AI names you when customers ask who’s best.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first page.