A sitemap is a list of the pages you want search engines to find, and it is only as useful as it is correct. One stray ampersand makes the whole file unreadable. An address with the wrong protocol, a date in the wrong format, or a page that robots.txt forbids is quietly ignored, and nothing tells you. This page checks the file against the sitemap protocol and Google's guidance, and lists what is wrong, where, and how to fix it.
Most validators stop at the protocol. This one also compares the sitemap with your robots.txt, because a sitemap that lists a page and a robots.txt that blocks it send opposite instructions. It checks the dates, the duplicates, the folder the sitemap covers, and the localized versions. Everything runs in your browser: the file is read on your device, and the code behind the page makes no network requests.
How to use the Sitemap Validator and Checker for XML Sitemaps
- Open or paste the sitemapOpen the file, including a .xml.gz file, or paste its text. If you only have the address, open it in a browser tab, save the page or copy its text, and use that.
- Read the summaryThe page says what kind of file it is, how many addresses it lists, how large it is, and how many problems there are. Problems to fix come first, then things to check, then notes.
- Fix what is listedEach finding says what is wrong, how many times it happens, and shows the first few places with their line numbers. A broken XML file shows the line and the text of the line.
- Add robots.txt and the addressOpen the settings and paste your robots.txt to find pages the sitemap lists but Googlebot may not crawl. Enter the address of the sitemap to check the folder rule.
- Download a tidier copyThe copy can drop duplicates, the tags Google ignores, and blocked addresses, and a file of more than 50,000 addresses is split into parts with an index.
What the protocol requires
The sitemap protocol at sitemaps.org needs three elements: a urlset that holds everything, a url for each page, and a loc with the page address. Everything else is optional. The file must be UTF-8. The five characters that mean something in XML have to be escaped in an address: an ampersand is written &, a single quote ', a double quote ", a greater than sign >, and a less than sign <. An address must be shorter than 2,048 characters.
A sitemap may hold at most 50,000 addresses and may be at most 50 MB, which is 52,428,800 bytes, when it is not compressed. A sitemap index has the same limits. All the addresses in a sitemap must be from a single host, and Google adds that, unless a sitemap is submitted through Search Console, it affects only the folder it is in and the folders below it. Placing it at the root of the site covers everything.
The root element must be in the namespace http://www.sitemaps.org/schemas/sitemap/0.9, written exactly like that: http rather than https, and with no slash at the end. A file with another namespace is well-formed XML but is not read as a sitemap, which is a very common reason for a sitemap that Search Console reports as having no addresses.
What Google says about the optional tags
Google's documentation is direct. It ignores the priority and changefreq values. It uses lastmod only when it is consistently and verifiably accurate, and the date should reflect a significant change to the page, not a minor one such as a new copyright year. Addresses should be fully qualified and absolute, because Google crawls them exactly as they are listed, and they should be the canonical addresses you want to appear in search.
The page uses these statements for its notes. A sitemap in which every page has the same lastmod is the common sign of a date that is the time the file was built. A priority and changefreq on every page only make the file larger. A lastmod in the future cannot be true. None of these is a protocol error, so they are shown as things to check, not things to fix.
Dates
The lastmod value uses the W3C date and time format. It can be a year, a year and month, a full date such as 2024-03-01, or a date and time with a time zone, such as 2024-03-01T10:30:00+01:00. The W3C note says a time is followed by a time zone designator, which is Z for UTC or an offset. The most useful forms are the full date and the full date with time, seconds, and zone.
The XML Schema that sitemaps.org publishes is stricter than the W3C format in a few places: it accepts only a full date or a date and time with seconds, so a year alone or a time without seconds is valid W3C but is rejected by strict validators. The page reports those as things to check, and reports as errors the values that no format allows, such as 2024-3-1, 03/01/2024, a space instead of the T, or a date that does not exist, like 30 February.
A sitemap that robots.txt contradicts
A sitemap tells crawlers which pages to visit, and robots.txt tells them which pages not to visit. When a page is in both, Googlebot follows robots.txt, so the page cannot be crawled. It may still appear in search results as an address with no description. This is easy to do by accident, for example when a staging rule or a rule for a search page is added to robots.txt and the sitemap is generated from a list of every page.
Paste your robots.txt into the settings and the page checks every address with the rules for Googlebot, using the same matching as Google's own parser: the most specific rule wins, an Allow wins a tie, and a group that names Googlebot replaces the group for all crawlers. Each blocked address is shown with the rule and the line of robots.txt that blocks it.
Duplicates and near duplicates
An address that appears twice is a plain mistake, and it is easy to find. The more common problem is one page under two addresses: with and without a trailing slash, with and without www, in http and https, or with a capital letter in the host. The page groups these and shows the two forms. Pick the one your site actually serves, usually the one it redirects to, and list only that.
Parameters that track a visitor, such as utm_source, gclid, or a session ID, produce a different address for the same page. They do not belong in a sitemap. The page also notes addresses that contain a fragment after #, a login, a broken % code, a space, or characters that should be percent-encoded.
Localized versions and large files
A sitemap can list the language and region versions of a page with xhtml:link entries. The page reads them and applies the same checks as the hreflang checker: valid codes, a link back from every version, and a link to itself, so a mistake in the localized entries is reported with the rest.
A file that is too large is split. The cleaned copy writes the addresses in parts of at most 50,000 and 50 MB, and a sitemap index that lists the parts, with the folder you give it. Each part keeps the namespaces and the other elements of the original, so images, video, and hreflang entries stay with their addresses.
How the checker was tested
The XML reader was compared with the Python expat parser on 841 documents: valid sitemaps and versions with a character deleted, a stray ampersand or bracket inserted, a tag renamed, an attribute repeated, or a prefix left undeclared. It gave the same verdict, and the same line for the first error, on all of them.
The protocol checks were compared with the XML Schemas that sitemaps.org publishes. Python's xmlschema validated 650 sitemaps and sitemap indexes, with errors in the address length, the dates, the change frequency, the priority, the order of the elements, the namespace, and the root, and the checker agreed with the schema on every one. Where the page is stricter or more lenient than the schema, for example with dates, the difference is written in this text.
Limits and accuracy
- The page cannot visit your site. Browsers do not let a page read another website, so it cannot check that each address works, redirects, or has a noindex tag. It checks the text of the file.
- It follows the sitemap protocol and Google's documentation. Other search engines may be stricter or more lenient.
- Extensions for images, video, and news are counted but not checked in detail. Localized versions are checked with the hreflang rules.
- A time of 24:00:00, which the XML Schema allows as the end of a day, is reported as an error.
- A file may be up to 120 MB after unpacking. Google's own limit is 50 MB, and the page reports a file larger than that.
- Only an XML sitemap can be cleaned and split. A sitemap index and a text sitemap are checked but not rewritten.
- The robots.txt check uses the rules for Googlebot. A group that names another crawler is not applied.
Frequently asked questions
How do I validate a sitemap?
Paste the sitemap or open the file here. The page checks that the XML is well-formed, that the namespace and elements follow the sitemap protocol, and that the addresses, dates, and limits are right. Each problem is listed with its line number.
Why does Search Console say my sitemap has no URLs?
The usual reasons are a wrong namespace, such as https instead of http in the xmlns value, XML that is not well-formed, so the file is rejected whole, or addresses that are relative paths. The page names each of these.
Does Google use priority and changefreq?
No. Google's documentation says it ignores both. It uses lastmod when it is consistently and verifiably accurate. You can leave the other two out, and the cleaned copy removes them.
What is the limit on the size of a sitemap?
A sitemap may hold at most 50,000 addresses and may be at most 50 MB, or 52,428,800 bytes, uncompressed. A larger site uses several sitemaps listed in a sitemap index, which has the same limits. The page can split a large file for you.
Can a sitemap list a page that robots.txt blocks?
It can, but it should not. Googlebot follows robots.txt, so it will not crawl the page, and the sitemap and the rules disagree. The page checks every address against your robots.txt and shows the rule that blocks it.
Which date format does lastmod use?
The W3C date and time format: a date such as 2024-03-01, or a date and time with a zone, such as 2024-03-01T10:30:00+01:00. A date written 03/01/2024 or 2024-3-1 is not valid.
Is my sitemap uploaded anywhere?
No. The file is read and checked in your browser, and the code behind the page makes no network requests. The page does not visit your site either, so what it checks is the text you give it.
Research and references
This page was written and checked against the sources below.

