SEO

Robots.txt Tester: Check If a URL Is Blocked

Robots.txt tester: paste your file, test URLs or a whole sitemap, see which rule decides, and check how Google, Bing, and AI crawlers are treated.

Use the Robots.txt Tester: Check If a URL Is Blocked

Open your robots.txt address in a browser tab, copy what it shows, and paste it here. A web page cannot fetch another site's file.

Leave it empty to test the home page. Paths such as /about work as well as full addresses.

Settings: the site this file belongs to

A robots.txt file covers only the host, protocol, and port it is served from. With a site entered, addresses on other sites or with another protocol are marked “Other site” instead of being judged by this file.

Paste a robots.txt file to see which URLs it blocks, which crawlers it affects, and what is wrong with it.

Everything runs in your browser. What you enter is never uploaded or stored.

A robots.txt file tells crawlers which parts of a site they may fetch, and a small mistake in it can hide a whole site from search or leave private paths open to every crawler. The rules look simple, but they are not applied top to bottom the way most people expect. The most specific rule wins, a crawler with its own group ignores the general group, and some crawlers ignore parts of the file altogether. This page tests a file against real URLs and shows which rule decided each one.

Paste your robots.txt, then a URL, a list of URLs, or a sitemap, and choose the crawler to test as. You see whether each address is allowed or blocked, the exact line that decided it, and every other rule that matched. A second view shows how every crawler in the list is treated, including the AI crawlers, and a third checks the file for mistakes. The page cannot fetch your live file, because browsers do not let one site read another's, so you open it in a tab and paste it. Everything runs in your browser, so what you paste or open is not uploaded.

How to use the Robots.txt Tester: Check If a URL Is Blocked

  1. Paste the robots.txt fileOpen the address of your file, such as your site's address followed by /robots.txt, copy what it shows, and paste it into the first box. You can also open a saved file. To test a change before you publish it, paste the new version.
  2. Enter the URLs to testType or paste one URL per line. Full addresses and plain paths such as /private/report.pdf both work. You can paste a whole sitemap and every address in it is tested. If you leave the box empty, the home page is tested.
  3. Choose the crawlerPick the crawler to test as, such as Googlebot or GPTBot, or type the name of another one. The file's rules for that crawler are applied the way the crawler reads them.
  4. Read the result and the deciding ruleEach URL shows Allowed or Blocked. Open a row to see which rule decided, which other rules matched, and why the others lost. Filter the list to blocked URLs only, or download the results as a CSV file.
  5. Check all crawlers and the fileThe All crawlers view shows how every crawler in the list treats one URL, which is the quick way to see whether the AI crawlers are blocked. The File check view lists mistakes with their line numbers.

How a crawler reads a robots.txt file

A robots.txt file is a list of groups. A group starts with one or more User-agent lines that name the crawlers it is for, and it continues with Allow and Disallow lines that give paths. A crawler looks for the group that names it and obeys only that group. If no group names it, it obeys the group for *, which means every other crawler. If neither exists, nothing is restricted. This is the behavior RFC 9309, the standard for robots.txt, describes.

The standard says crawlers match the product token without regard to case, and that when several groups name the same crawler their rules are combined into one. So two groups for Googlebot, even far apart in the file, act as one. A crawler that has its own group does not also read the * group. A common mistake is to write a Disallow rule under * and then add a short group for a particular crawler, expecting the general rules to still apply to it. They do not.

A file only covers the host, protocol, and port it is served from. Google states that the rules of https://example.com/robots.txt do not apply to http://example.com, to a subdomain, or to a different port. Each needs its own file.

Which rule wins

When several rules in a group match a URL, the most specific one is used. RFC 9309 defines that as the match with the most characters, and Google describes it as the rule with the longest path. Order in the file does not matter. If an Allow rule and a Disallow rule match with the same length, the standard says the Allow rule should be used, and Google says it uses the least restrictive rule.

That is why a short Allow can open up a hole in a long Disallow, and the reverse. With Disallow: /folder and Allow: /folder/, the URL /folder/page is allowed, because the Allow rule is one character longer. With Allow: /page and Disallow: /*.htm, the URL /page.htm is blocked, because the wildcard rule is longer. Click a row in the results to see the matching rules in order of length.

  • * matches any run of characters, including none.
  • $ at the end of a path means the URL must end there, so /*.pdf$ blocks PDF files but not /file.pdf?download=1.
  • A path has to start with / or *. A path such as admin never matches anything.
  • Paths are case sensitive, so Disallow: /Private does not block /private.
  • The query string is part of what is matched, so /*? blocks every URL with a query.
  • An empty Disallow: line blocks nothing.
  • The robots.txt file itself can always be fetched.

Crawlers that do not follow the general group

Google lists several special-case crawlers that follow robots.txt only in a limited way. AdsBot-Google and Mediapartners-Google, the crawler AdSense uses to read pages so it can choose relevant ads, ignore the * group. Google says they follow only a group that names them. A site that blocks everything with User-agent: * and Disallow: / therefore does not block these crawlers, and a site that wants to block AdSense has to say so by name.

Several Google crawlers fall back to Googlebot. Googlebot-Image, Googlebot-News, and Googlebot-Video use their own group if there is one, and the Googlebot group if not. Google-InspectionTool, which powers Google's testing tools, falls back the same way. This tool applies those fallbacks when you test as one of them.

A few tokens are not crawlers at all. Google-Extended and Applebot-Extended are control tokens. They crawl nothing. They tell the company whether content it has already crawled may be used for its generative AI models. Google says Google-Extended does not affect inclusion in Google Search, and Apple says pages that disallow Applebot-Extended can still appear in search results.

AI crawlers and robots.txt

Since 2023 many site owners have used robots.txt to say whether AI companies may use their content. The companies publish the names of their crawlers, and the list on this page follows their documentation. OpenAI says GPTBot crawls content that may be used to train its generative AI models, and OAI-SearchBot decides whether a site is shown in ChatGPT search answers. Anthropic documents ClaudeBot for training data, Claude-SearchBot for search quality, and Claude-User for fetches a person triggers. Perplexity documents PerplexityBot for its search results. Common Crawl documents CCBot, and Meta documents Meta-ExternalAgent.

The distinction between crawlers that gather data and fetchers that act for a person matters. OpenAI says robots.txt rules may not apply to ChatGPT-User because a user starts the action. Perplexity says Perplexity-User generally ignores robots.txt, and Meta says Meta-ExternalFetcher may bypass it. A Disallow rule for these is a request that the operator says it may not honor, so the tool marks them as such. Blocking a crawler for training and allowing the search crawler of the same company is a common choice, and the All crawlers view shows both side by side.

Whether to block these crawlers is a decision for you. Blocking a training crawler stops that company from collecting new content from the site. Blocking a search crawler can keep your pages out of that company's answers. The tool does not recommend either. It shows what your file does today.

What robots.txt cannot do

Google says robots.txt is mainly for managing crawler traffic and is not a way to keep a page out of Google Search. A page that is disallowed can still be indexed if other sites link to it, and then the address and perhaps the anchor text of the links can appear in results without a description. To keep a page out of search results, leave it crawlable and use a noindex meta tag or response header, or put it behind a password. If the page is blocked, the crawler can never see the noindex tag.

Robots.txt is also not security. It is a public file, and listing a private path in it tells everyone where the path is. Well-behaved crawlers follow it, but nothing forces a crawler to. Protect private content with authentication.

Google reads only four fields: user-agent, allow, disallow, and sitemap. It does not support Crawl-delay, and noindex or nofollow lines have no effect. Some other crawlers do support Crawl-delay, for example Anthropic documents it for ClaudeBot. The File check view marks lines that most crawlers ignore.

Size, errors, and how often it is read

Google and the standard read at most 500 KiB of the file and ignore the rest, so rules after that point do nothing. The File check view finds the line where the limit falls. The standard also says crawlers should not use a cached copy for more than 24 hours unless the file is unreachable, so a change can take a day to take effect.

What the server answers for /robots.txt matters as much as what is in it. Under the standard, a 4xx answer means the crawler may fetch anything, and a 5xx answer means it must assume everything is disallowed. Google treats most 4xx answers as no restrictions, and when the server errors it pauses crawling and later uses its last copy for up to 30 days. It follows at least five redirects and then treats the file as missing. A robots.txt that is accidentally served with a server error can stop crawling of the whole site, even though its content is fine.

Mistakes the file check looks for

The check reads the file the way Google's open-source parser does, including its tolerance for a few misspellings, and reports what it finds. Google's parser accepts forms such as useragent, dissallow, and disalow, but a crawler that follows the standard strictly may ignore those lines, so the check asks you to spell them correctly.

  • Rules before the first User-agent line, which belong to no group and are ignored.
  • Misspelled directives, a missing colon, and lines that cannot be understood.
  • A * group that blocks the whole site.
  • Rules that block CSS or JavaScript files that Google needs to display a page, unless an Allow rule opens them again.
  • A Disallow rule for the AdSense crawler.
  • Sitemap lines that are not full addresses.
  • Paths that do not start with /, and $ in the middle of a path.
  • Repeated rules, Allow and Disallow rules for the same path, and groups for the same crawler that are combined.
  • A file larger than 500 KiB.

Limits and accuracy

  • The page cannot fetch your live robots.txt, because browsers do not let a page read another site's files. Paste the file or open a saved copy.
  • The rules are those of RFC 9309 and Google's documentation: the longest match wins, Allow wins a tie, and * and $ are wildcards. Not every crawler follows them exactly, so the result is what the standard says, not a promise about one particular crawler.
  • It cannot know what your server actually answers for /robots.txt, including status codes and redirects, which change what crawlers do. Check those in your server logs or Search Console.
  • Crawler names and descriptions come from each company's public documentation and can change. Check the company's page if a decision depends on it.
  • It does not verify who is making a request. Anyone can claim a crawler's name, and the tool only tests the file.
  • Up to 5,000 URLs are tested at once, and a file or list may be up to 5 MB.

Frequently asked questions

How do I test whether a URL is blocked by robots.txt?

Paste the contents of the robots.txt file, enter the URL, and choose the crawler. The tool shows whether the URL is allowed or blocked and which line of the file decided it. Paste a sitemap instead of one URL to test every page in it at once.

If two rules conflict, which one applies?

The longer, more specific rule applies, and the order in the file does not matter. If an Allow and a Disallow rule match with the same length, the Allow rule wins. Open a row in the results to see all the matching rules and which one decided.

Does Disallow remove a page from Google?

No. Google says robots.txt is not a way to keep a page out of search results, and a disallowed page can still be indexed if other sites link to it. To keep a page out of results, let it be crawled and add a noindex meta tag or header, or protect it with a password.

Why did Disallow: / not stop the AdSense crawler?

Google says its special-case crawlers, including AdsBot-Google and Mediapartners-Google, ignore the * group and follow only a group that names them. A rule under User-agent: * does not reach them. To block one, write a group for it by name.

Should I block GPTBot, ClaudeBot, and other AI crawlers?

That is your decision, and the tool does not make it for you. Blocking a training crawler stops that company from collecting new content from your site. Blocking a search crawler may keep your pages out of that company's answers. Some user-triggered fetchers are documented as possibly ignoring robots.txt.

Why does the tool not fetch my robots.txt by itself?

A web page cannot read a file from another site, because browsers block it for security. Fetching it would need a server to do the request, which would mean sending your address to a server. Pasting the file keeps everything on your device.

Is my robots.txt uploaded when I use this page?

No. The tests run in your browser and the code behind the page makes no network requests, so the file and the URLs you paste stay on your device. You can check this in the Network panel of your browser's developer tools.

Research and references

This page was written and checked against the sources below.

  1. RFC 9309: Robots Exclusion Protocol
  2. Google Search Central: How Google interprets the robots.txt specification
  3. Google Search Central: Introduction to robots.txt
  4. Google Search Central: Google's common crawlers
  5. Google Search Central: Google's special-case crawlers
  6. OpenAI: Overview of OpenAI crawlers
  7. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
  8. Google's open-source robots.txt parser (GitHub)