SEO

Hreflang Generator and Checker for Multilingual Sites

Hreflang generator and checker: build tags, a header, or a sitemap, then check that every page links back. Free, and nothing is uploaded.

Use the Hreflang Generator and Checker for Multilingual Sites

Add one line for each version of the page: its language code and its full address. Put the same set on every version.

Add the language codes and addresses of your page versions to get the tags.

Everything runs in your browser. What you enter is never uploaded or stored.

A site that has the same page in several languages or for several countries uses hreflang to tell Google which version is meant for which visitor. The tags are short, and the rules behind them are strict: every version has to list all the others and itself, the codes have to come from particular lists, and the addresses have to be complete. A single wrong code or a missing return link and Google ignores the annotation. Most of the mistakes cannot be seen by looking at the page.

This page does both jobs. Add the language codes and addresses of your pages and it writes the tags, the header, or a sitemap. Or paste what you already have and it checks it, and lists every problem with the page it is on and the tag that would fix it. Nothing is uploaded: the text is read in your browser, and the code behind the page makes no network requests.

How to use the Hreflang Generator and Checker for Multilingual Sites

  1. Choose generate or checkUse the first tab to write new annotations, or the second to check ones you already have.
  2. Add the versions of your pageFor one page, type a language code and a full address on each line. For many pages, paste a table with the codes in the first row and one row for each page. Each code is checked as you type, with a suggestion if it is wrong.
  3. Choose x-defaultPick the address for visitors whose language matches none of the versions, such as your main version or a language selector. It is optional.
  4. Copy the resultPick HTML tags, an HTTP header, or an XML sitemap, and copy or download it. The page checks what it wrote and tells you if anything is wrong with the set.
  5. Check what is liveTo check tags you already use, paste them, a header, or a sitemap into the second tab. If the text is the tags of one page, add the address of that page so it can check that the page lists itself.

What hreflang does and how it is written

Hreflang tells a search engine that a group of pages are the same page for different audiences, and which audience each is for. Google's documentation describes three ways to give it: link elements in the head of the page, an HTTP header, and a sitemap that lists every version with xhtml:link entries. According to Google the three methods are equivalent, and using more than one has no benefit in Search and makes the setup harder to manage, so choose one.

The HTML form is a link element with rel="alternate", an hreflang value, and the address of that version. It goes on every version of the page. The header form carries the same values in a Link header and is meant for files with no head, such as PDFs. The sitemap form puts the whole set under each address and suits a large site, because there is nothing to change in the pages themselves.

The rules that decide whether it works

Google states the central rule plainly: each language version must list itself as well as all the other language versions, and if two pages do not both point to each other, the tags are ignored. That means the set has to be written the same way on every page, and it is the reason that a generated set is safe and a hand-edited one often is not. Adding a new language to one page but not to the others breaks the links for the new page.

Addresses must be fully qualified, with the protocol, such as https://example.com/de/. A path such as /de/ is not valid. The address in the tag also has to match the real address of the page, because a page that redirects from one form to another, such as with and without a trailing slash, is not the page the tag names. The checker compares the addresses exactly and tells you when two of them differ only by a trailing slash, http against https, or www.

The x-default value names the page for visitors whose language matches none of the versions. Google recommends a fallback page for unmatched languages, especially on language and country selectors or on home pages that redirect automatically, and says x-default was designed for language selector pages.

Language, script, and region codes

The value is a language code from ISO 639-1, optionally followed by a region code from ISO 3166-1 alpha-2, such as en-GB for English in the United Kingdom. Google's documentation also allows a script from ISO 15924, such as zh-Hant for Traditional Chinese. Codes are separated with a hyphen, never an underscore. A code that is a region alone is not valid: Google's example is that be, for Belgium, does not work, and de-BE, nl-BE, or fr-BE should be used instead.

The mistakes that come up most are listed below. The checker finds each of them and shows the value to use instead.

  • en-UK. The United Kingdom is GB in ISO 3166-1, so the value is en-GB.
  • es-419 and other UN region numbers. Google's documentation names ISO 3166-1 alpha-2 codes only, and lists es-419 as unsupported.
  • en_US with an underscore, which should be en-US.
  • A country written alone, such as us or gb, which is a region and not a language.
  • iw, in, ji, jw, and mo, old language codes that were replaced by he, id, yi, jv, and ro.
  • Three-letter codes such as eng or fil. Google documents two-letter codes, and the checker shows the two-letter form when there is one.
  • The wrong order, such as zh-CN-Hans. The order is language, script, then region.

Many pages: a table instead of a form

For a site with hundreds of pages, typing each version is not practical. In the second mode you paste a table, copied from a spreadsheet or opened from a CSV file. The first row holds the language codes, and each following row is one page, with the address of that page in each language in the cell under its code. A column named x-default holds the fallback address. A page that has no version in some language leaves that cell empty.

The page turns each row into a set in which every address lists the whole row, so the self reference and the return links are right by construction, and writes one sitemap for all of them. You can also take the HTML tags for any one page from the list. Because the sitemap is written from the table, a change in the table and a new download is all it takes to add a language.

How the check works across a set of pages

Checking one page's tags can only show that the codes and addresses are well formed. The return links need the other pages. When you paste a sitemap, which lists many pages with their alternates, the checker can follow every link: for each page A that points to page B, it looks for B in the sitemap and checks that B points back to A. If it does not, the finding names the page that lacks the tag and gives the tag to add, with the code that page A uses for itself.

It also compares the pages of a set. If most pages list a French version and one does not, that page is named. Where an alternate address is not in the text you pasted, the checker cannot see the page, so it says that the return link could not be checked instead of calling it an error.

How the checker was tested

The language, region, and script lists come from the pycountry package: 184 two-letter languages, 249 regions, and 226 scripts. The test for the codes compares the verdict of the page with an independent implementation of Google's documented format for 434 codes, valid and invalid, written in different cases. The reading of annotations was tested on 120 texts written by a separate Python program in HTML, header, sitemap, and line form, with random quotes, attribute order, and entities.

To test the checker, a program built 150 sitemaps, each with one known mistake injected: an invalid code, a relative address, a page that does not list itself, a missing return link, two addresses for one code, a repeated tag, mixed protocols, two x-default values, and a region alone. The checker found the mistake by name in every case, and found no error in the sets that had none.

Limits and accuracy

  • The page cannot fetch your pages. Browsers do not let a page read another site, so it checks the text you paste, not your live site. A tag that is correct here can still fail if the real page redirects, returns an error, or has a canonical tag that points to another address.
  • Return links can be checked only between pages that are in the text. For the tags of a single page, the page checks the codes and addresses and, if you give the address of the page, the self reference.
  • It follows Google's documented format. Other search engines have their own rules, and Google can still choose not to use an annotation that is valid.
  • It checks that a code is well formed and exists. It cannot tell whether the content of a page is really in that language.
  • A canonical tag is checked only when it is in the text you paste. Paste the whole head of a page to include it.
  • A sitemap of up to 20 MB can be opened. Google's own limit for a sitemap is 50,000 addresses and 50 MB, and the checker reports more than 50,000 addresses.
  • The tables cover the codes of ISO 639-1, ISO 3166-1, and ISO 15924 as they were when this page was built. A code added later is not known to it.

Frequently asked questions

What is hreflang?

It is an annotation that tells a search engine that a group of pages are the same page for different languages or countries, and which version is meant for which audience. It is written as link tags in the head of each page, as an HTTP header, or in a sitemap.

How do I generate hreflang tags?

Add a line for each version of the page with its language code, such as en-GB, and its full address. Choose an x-default if you have one. The page writes the tags, the header, or a sitemap, and checks that the set links to itself. Put the same set on every version.

Why does Google ignore my hreflang?

The most common reason is a missing return link: if page A points to page B, page B must point back, or the tags are ignored. Others are a wrong code such as en-UK, a relative address, a page that does not list itself, or two addresses for the same code. The checker looks for all of these.

Is en-UK a valid hreflang value?

No. The region code for the United Kingdom is GB, so the value is en-GB. Google's documentation names ISO 3166-1 alpha-2 region codes, and UK is not one of them.

Do I need x-default?

It is optional. Google recommends a fallback page for visitors whose language matches none of your versions, and says x-default was designed for language selector pages. If you have no such page, you can leave it out.

Should I use HTML tags, a header, or a sitemap?

Google says the three are equivalent and that there is no benefit in using more than one. Tags suit a small site, a header suits files with no HTML such as PDFs, and a sitemap suits a large site because the pages themselves do not change.

Is my site uploaded when I use this page?

No. The addresses and tags are read and written in your browser, and the code behind the page makes no network requests. The page does not visit your site either, so what it checks is the text you give it.

Research and references

This page was written and checked against the sources below.

  1. Google Search Central: Tell Google about localized versions of your page
  2. Library of Congress: ISO 639-2 and 639-1 language code list
  3. Unicode: ISO 15924 four-letter script codes
  4. ISO 3166 country codes