SEO

Robots.txt Generator: Rules, AI Crawlers, and a Live Test

Robots.txt generator in your browser: build rules for Google, Bing, and AI crawlers, get warnings before you publish, and test any path against the result.

Use the Robots.txt Generator: Rules, AI Crawlers, and a Live Test

Start from

A starting point adds its rules to what you have. Hover or focus a button for what it does.

Rules for crawlers

Block crawlers from the whole site

AI training
AI search
Fetches for a user

Hover or focus a name for what its operator says it does. Some fetchers that act for a person may not follow robots.txt at all.

Your robots.txt

Try an address

AllowedNo rule in the group for * matches this path.

Everything runs in your browser. What you enter is never uploaded or stored.

A robots.txt file is short, and a short mistake in it can keep a whole site out of search, or leave a private folder open to every crawler. The common generators fill in a form and give you text, but they do not tell you what the text will do. They do not warn that a crawler with its own group ignores the general group, that a rule for CSS and JavaScript can break how Google sees your pages, or that Disallow: / for everyone does not reach the AdSense crawler.

This page builds the file from the rules you give it, and then shows what the file does. Choose a starting point, add the groups and paths that you want, tick the AI crawlers that you want to keep out, and read the warnings that appear. A last box tests any path against the finished file as Google, Bing, or one of the AI crawlers, and names the rule that decided. Everything is done in your browser, and the code behind the page makes no network requests, so a staging site's paths are not sent anywhere.

How to use the Robots.txt Generator: Rules, AI Crawlers, and a Live Test

  1. Choose a starting point, or begin emptyWordPress keeps crawlers out of the admin area and leaves admin-ajax.php open. Online shop keeps them out of the cart, the checkout, and search results. Block AI training crawlers adds one group for the crawlers whose operators say they train models. Staging blocks everything, and it is not security. You can combine them, and then edit the rules.
  2. Write the groupsA group names the crawlers it is for, and lists the paths that they may not fetch and the paths that they may. Use * for every crawler that has no group of its own. A path starts with a slash. A * inside a path matches any text, and a $ at the end means that the address must end there.
  3. Tick the crawlers to blockThe crawlers are listed by what they are for. Blocking a crawler that gathers training data keeps new content from that company. Blocking a search crawler can keep your pages out of its answers. Hover a name to read what its operator says about it.
  4. Read the warnings and try some addressesThe notes under the file say what is wrong or risky, before you publish. Then choose a crawler and write a path in the last box to see whether it is allowed, and by which rule. When it is right, copy it or download robots.txt and put it at the root of your site.

What the file is, and where it goes

A robots.txt file is a plain text file at the root of a site, at the address of the site followed by /robots.txt. It lists groups. A group starts with one or more User-agent lines that name the crawlers it is for, and it continues with Allow and Disallow lines that give paths. The Sitemap line is not part of a group. It names a sitemap, with a full address, and it can be anywhere in the file. The rules are set out in RFC 9309, the standard for the file, and Google describes how it reads them.

A file covers only the host, protocol, and port that it is served from. The file at https://example.com/robots.txt does not apply to http://example.com, to a subdomain such as blog.example.com, or to another port, and each of them needs its own. The file has to be encoded in UTF-8 and is read up to 500 KiB. This page writes only plain ASCII lines, so the encoding is not a risk.

Which rule wins

When several rules in a group match an address, the standard says that the most specific one is used, which means the one with the longest path. The order of the lines does not matter. If an Allow and a Disallow rule match with the same length, the Allow rule is used. So Disallow: /wp-admin/ with Allow: /wp-admin/admin-ajax.php blocks the admin area and leaves the one file open, because the Allow rule is longer for that address.

The page writes the Allow lines before the Disallow lines of a group. The order makes no difference to a crawler that follows the standard, but some older parsers use the first rule that matches, and for them an Allow that comes first behaves the way the writer meant.

A crawler with its own group ignores the * group

This is the mistake that this page is built to prevent. A crawler looks for a group that names it and obeys only that group. If there is none, it obeys the group for *. If you write Disallow: /private/ for * and then add a group for Googlebot that only has Disallow: /drafts/, Googlebot is no longer held to the /private/ rule, because it has a group of its own. Many files are wrong in this way, and nobody notices, because the file is valid.

The option to repeat the rules of * in the other groups does it for you. When a group has rules of its own, the rules of the * group are added to it in the file, so that the crawler named in it is held to both. A group that blocks everything is left as it is, because it has nothing to add. If you turn the option off, the page warns you that those crawlers are not held to the * rules.

AI crawlers

The operators of AI products publish the names of their crawlers, and the list on this page follows their documentation. OpenAI documents GPTBot, which it says crawls content that may be used to train its models, OAI-SearchBot, which decides whether a site is shown in ChatGPT search answers, and ChatGPT-User, which visits pages when a person asks. Anthropic documents ClaudeBot, Claude-SearchBot, and Claude-User in the same way. Google-Extended and Applebot-Extended are not crawlers. They are tokens that tell Google and Apple whether content they have crawled may be used for their AI models, and they do not affect search. Common Crawl documents CCBot, Perplexity documents PerplexityBot, and Meta documents Meta-ExternalAgent.

Whether to block them is your choice, and the page does not make it. A training crawler that is blocked can no longer collect new content from the site. A search crawler that is blocked can keep your pages out of that company's answers. Some fetchers that act for a person are documented as possibly ignoring robots.txt, such as ChatGPT-User, Perplexity-User, and Meta-ExternalFetcher, so a rule for them is a request that the operator says it may not honor. The list says so when you hover a name.

What the warnings mean

A path has to start with a slash, or with a * that stands for the start. A path such as private matches nothing, and the page says so. A space in a path is an error, and has to be written as %20. A rule that blocks folders that usually hold CSS and JavaScript, such as /wp-includes/ or /assets/, or any address that ends in .css or .js, gets a warning, because Google asks that these files are not blocked, since it needs them to render a page as a visitor sees it.

A group for * that disallows the whole site gets a warning of its own. It is right for a staging site, but it is not security, because it asks crawlers to stay away and does not stop them, and Google says that its AdSense and ad quality crawlers ignore the * group. A site that has to be private needs a password. A Crawl-delay line gets a note, because Google does not read it, while Bing and some others do. A sitemap that is not a full address, with https:// and the host, is an error, because a crawler does not guess the host.

What robots.txt cannot do

Google says that robots.txt is meant for managing crawler traffic, and that it is not a way to keep a page out of Google Search. A page that is disallowed can still be indexed if other sites link to it, and then it can appear without a description. To keep a page out of results, let it be crawled and use a noindex tag or header, or put it behind a password. If a page is blocked in robots.txt, the crawler can never see the noindex tag. The file is also public, and a list of private paths in it tells everyone where they are.

How it was tested

The file that the page writes was tested with Protego, a robots.txt parser for Python that follows the rules of Google. 300 sets of rules were made at random, with a general group, groups for single crawlers, and groups that name several, with allowed and disallowed paths, with and without the repeated rules. The page wrote the file for each, and Protego was asked 40,500 questions about it: whether each of nine crawlers may fetch each of 15 paths. The page's own tester and Protego gave the same answer to every one. A small model of the settings, written for the test and not using the page's matcher, gave the same answer to every one as well, which shows that the file means what the settings say.

Every file was also read by the tester on this site without an error. The warnings were tested on written examples. The page was not tried against live crawlers, because that cannot be done from a browser, and nobody can promise what a crawler does with a file.

Limits and accuracy

  • The page builds a file and tests it against the rules of RFC 9309 and Google's description of them. A crawler can read it in its own way, and some ignore it altogether. A file is a request, and it does not protect anything.
  • It does not fetch your live file, and cannot check what your server answers for /robots.txt, including status codes and redirects, which change what crawlers do.
  • The crawlers and their purposes come from their operators' public documentation, which was read for the tester on this site and can change. Check the operator's page if a decision depends on it.
  • The groups that the page writes name each crawler by the name that its operator gives. It does not verify who makes a request, because anyone can use a crawler's name.
  • Only the four fields that Google reads are written: User-agent, Allow, Disallow, and Sitemap, with Crawl-delay if you ask. Directives that crawlers ignore, such as noindex, are not offered.
  • The warnings are about common mistakes. A file that has none can still block something that you want crawled, so try the addresses that matter to you.

Frequently asked questions

How do I create a robots.txt file?

Choose a starting point or begin empty, write the paths that crawlers may not fetch, tick any AI crawlers to block, and add your sitemap address. Copy the file, or download it, and put it at the root of your site, so that it opens at your address followed by /robots.txt.

Where does robots.txt go?

At the root of the host, as /robots.txt. A file in a folder is not read. Each host, protocol, and port needs its own, so a subdomain such as blog.example.com has to have a file of its own.

How do I block AI crawlers like GPTBot and ClaudeBot?

Tick them in the list, or use the starting point that blocks AI training crawlers. The page writes one group with a User-agent line for each, and Disallow: /. Search crawlers are not affected. Some fetchers that act for a person may not follow robots.txt at all.

Why does my crawler ignore the * rules?

Because a crawler that has a group of its own obeys only that group. The option to repeat the * rules in the other groups adds them to each group that has rules, so that the crawler is held to both.

Does Disallow keep a page out of Google?

No. Google says that robots.txt is not a way to keep a page out of search results. A disallowed page can still be indexed if other sites link to it. Use a noindex tag or header on a page that can be crawled, or a password.

Should I block CSS and JavaScript?

No. Google needs them to render your pages the way visitors see them, and it asks that they are not blocked. The page warns about folders and extensions that usually hold them.

Does Google read Crawl-delay?

No. Bing and some other crawlers do. The page writes it if you ask, with a note that Google does not read it.

Is anything I type uploaded?

No. The file is built and tested in your browser, and the code behind the page makes no network requests. Nothing is saved, so copy or download the file before you close the page.

Research and references

This page was written and checked against the sources below.

  1. RFC 9309: Robots Exclusion Protocol
  2. Google Search Central: How Google interprets the robots.txt specification
  3. Google Search Central: Create a robots.txt file
  4. Google Search Central: Google's special-case crawlers
  5. OpenAI: Overview of OpenAI crawlers
  6. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
  7. Protego: a robots.txt parser (GitHub)