XML Sitemap Validator
Every protocol rule checked. No URL limit, no signup.
Everything runs in this tab. Nothing you paste is uploaded, logged or sent anywhere. Open your network panel and check.
Paste your sitemap above, or enter its address and press Fetch, and every rule in the sitemaps.org protocol is checked at once: the URL count against the 50,000 limit, the uncompressed size against the 50 MB limit, and the format of every loc, lastmod, changefreq and priority. Sitemap index files get their own rule set.
Most people arrive from Search Console holding "Sitemap could not be read" or "Sitemap can be read, but has errors", neither of which names a line. Each finding here quotes the value it objected to, cites the rule, and says what to write instead.
Two things separate it from the rest of this search. Nothing is uploaded, so a staging sitemap listing pages you have not launched stays on your machine. And checks that could not run are reported as not run rather than as passed.
The two hard limits
A single sitemap may list 50,000 URLs and may be 50 MB uncompressed, which the protocol states as 52,428,800 bytes. Both are checked, and the byte size is reported whether or not you are close to it, because it is the one number you cannot judge by looking at the file.
Gzip buys no headroom: the limit applies after decompression, so a 3 MB sitemap.xml.gz that expands to 61 MB is over it. A sitemap index carries the same two limits. Both warn at ninety per cent, since a sitemap built from a growing catalogue crosses the line on an ordinary Tuesday.
- 50,000 <url> entries per sitemap file.
- 52,428,800 bytes uncompressed, whether or not you gzip it.
- 2,048 characters per <loc> value.
- 50,000 <sitemap> entries per sitemap index.
The namespace has to be exact
The root is <urlset>, or <sitemapindex> for an index, and it must declare http://www.sitemaps.org/schemas/sitemap/0.9. Three details catch people out: it is http and not https even on an https site, it has no trailing slash, and the version is 0.9 and always has been.
Get it wrong and nothing complains loudly. The file is still well-formed and still opens in a browser. Search engines simply do not recognise elements in the wrong namespace, so they read your sitemap as a document containing nothing.
Mixing <url> and <sitemap> children is an error, because a file is either a sitemap or an index. The declared encoding must be UTF-8. Extension elements are checked only for having a declared prefix; the image, video and news rules themselves are not validated here.
lastmod, changefreq and priority
lastmod must be a W3C Datetime. YYYY-MM-DD is valid on its own; a full timestamp is valid if it carries a zone, as in 2026-03-14T09:30:00+00:00. A space instead of the T, a Unix timestamp, 14/03/2026, or a timestamp with no zone are all errors.
A second lastmod check runs that no other free validator does. If more than four entries carry one and every value is identical, you get a warning: your generator is stamping build time rather than the date content changed. Google uses lastmod only when it is consistently and verifiably accurate, so that pattern teaches it to ignore the field.
changefreq accepts seven lowercase values: always, hourly, daily, weekly, monthly, yearly, never. priority must be between 0.0 and 1.0. Both are also reported as information when present at all, because Google states in its own documentation that it ignores them.
Same host, same directory, and fetching by URL
A sitemap at https://example.com/catalog/sitemap.xml may list URLs under https://example.com/catalog/ and may not list URLs under https://example.com/images/. Same host, same scheme, at or below its own directory. Those rules need the address the sitemap will be published at, which the file does not contain.
That is what the URL box is for, and you can type an address without fetching. Enter it and three checks switch on: another host is an error, a different scheme is an error, and a loc outside the sitemap own directory is a warning. Leave it empty and all three are listed as not run rather than passing quietly.
Fetch requests the file from your browser straight to your server with no proxy in between, so CORS blocks it on many hosts. When it does you get the reason and a curl command to run yourself.
Checking a sitemap in code
Worth wiring into a build step, so a sitemap that has outgrown the 50,000 URL limit fails your deploy rather than your indexing. Each sample counts URLs, measures the uncompressed size and checks lastmod.
// Node 18 or later. npm i @xmldom/xmldom
import { DOMParser } from '@xmldom/xmldom';
const SITEMAP_NS = 'http://www.sitemaps.org/schemas/sitemap/0.9';
const MAX_URLS = 50_000;
const MAX_BYTES = 52_428_800;
const response = await fetch('https://example.com/sitemap.xml');
const xml = await response.text();
const bytes = Buffer.byteLength(xml, 'utf8');
// xmldom does not resolve external entities, so there is no XXE surface here.
const doc = new DOMParser().parseFromString(xml, 'text/xml');
const root = doc.documentElement;
if (root.namespaceURI !== SITEMAP_NS) {
throw new Error('Namespace is "' + root.namespaceURI + '", expected "' + SITEMAP_NS + '"');
}
const locs = doc.getElementsByTagNameNS(SITEMAP_NS, 'loc');
console.log(locs.length + ' URLs, ' + bytes + ' bytes uncompressed');
if (locs.length > MAX_URLS) throw new Error('Over the 50,000 URL limit');
if (bytes > MAX_BYTES) throw new Error('Over the 50 MB limit');
const lastmods = doc.getElementsByTagNameNS(SITEMAP_NS, 'lastmod');
for (let i = 0; i < lastmods.length; i++) {
const v = lastmods.item(i).textContent.trim();
// W3C Datetime: a bare date, or a timestamp that carries a zone.
const ok = /^\d{4}-\d{2}-\d{2}$/.test(v)
|| (/^\d{4}-\d{2}-\d{2}T/.test(v) && /(Z|[+-]\d{2}:\d{2})$/.test(v));
if (!ok) console.error('Invalid lastmod: ' + v);
}# pip install lxml requests
import re
import sys
import requests
from lxml import etree
SITEMAP_NS = 'http://www.sitemaps.org/schemas/sitemap/0.9'
MAX_URLS, MAX_BYTES = 50_000, 52_428_800
raw = requests.get('https://example.com/sitemap.xml', timeout=30).content
# The three flags are the safe form. lxml will otherwise expand entities and
# fetch a DTD the document references.
parser = etree.XMLParser(resolve_entities=False, no_network=True, load_dtd=False)
root = etree.fromstring(raw, parser)
if etree.QName(root).namespace != SITEMAP_NS:
sys.exit('Wrong namespace: %s' % etree.QName(root).namespace)
locs = root.findall('{%s}url/{%s}loc' % (SITEMAP_NS, SITEMAP_NS))
print('%d URLs, %d bytes uncompressed' % (len(locs), len(raw)))
if len(locs) > MAX_URLS:
sys.exit('Over the 50,000 URL limit')
if len(raw) > MAX_BYTES:
sys.exit('Over the 50 MB limit')
W3C = re.compile(r'^\d{4}-\d{2}-\d{2}(T\d{2}:\d{2}(:\d{2}(\.\d+)?)?(Z|[+-]\d{2}:\d{2}))?$')
for lm in root.iter('{%s}lastmod' % SITEMAP_NS):
if not W3C.match((lm.text or '').strip()):
print('Invalid lastmod:', lm.text)
# The protocol publishes an XSD, which lxml can check structure against:
# schema = etree.XMLSchema(etree.parse('sitemap.xsd'))
# schema.assertValid(root.getroottree())import javax.xml.XMLConstants;
import javax.xml.parsers.DocumentBuilderFactory;
import java.io.ByteArrayInputStream;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.LocalDate;
import java.time.OffsetDateTime;
import java.time.format.DateTimeParseException;
import org.w3c.dom.Document;
import org.w3c.dom.NodeList;
final String NS = "http://www.sitemaps.org/schemas/sitemap/0.9";
byte[] bytes = HttpClient.newHttpClient()
.send(HttpRequest.newBuilder(URI.create("https://example.com/sitemap.xml")).build(),
HttpResponse.BodyHandlers.ofByteArray())
.body();
DocumentBuilderFactory f = DocumentBuilderFactory.newInstance();
// Without setNamespaceAware(true) every namespace lookup below returns nothing.
// This is the single most common reason sitemap parsing code silently finds zero URLs.
f.setNamespaceAware(true);
f.setFeature(XMLConstants.FEATURE_SECURE_PROCESSING, true);
f.setFeature("http://apache.org/xml/features/disallow-doctype-decl", true);
f.setXIncludeAware(false);
Document doc = f.newDocumentBuilder().parse(new ByteArrayInputStream(bytes));
NodeList locs = doc.getElementsByTagNameNS(NS, "loc");
System.out.printf("%d URLs, %d bytes uncompressed%n", locs.getLength(), bytes.length);
if (locs.getLength() > 50_000) System.err.println("Over the 50,000 URL limit");
if (bytes.length > 52_428_800) System.err.println("Over the 50 MB limit");
NodeList lastmods = doc.getElementsByTagNameNS(NS, "lastmod");
for (int i = 0; i < lastmods.getLength(); i++) {
String v = lastmods.item(i).getTextContent().trim();
try {
if (v.length() == 10) LocalDate.parse(v);
else OffsetDateTime.parse(v); // rejects a timestamp with no zone
} catch (DateTimeParseException e) {
System.err.println("Invalid lastmod: " + v);
}
}using System.Globalization;
using System.Xml;
const string Ns = "http://www.sitemaps.org/schemas/sitemap/0.9";
// XmlReader streams, so a 50 MB sitemap never becomes a 400 MB DOM.
var settings = new XmlReaderSettings
{
DtdProcessing = DtdProcessing.Prohibit, // external entities are never fetched
XmlResolver = null,
IgnoreWhitespace = true,
};
string[] w3c =
{
"yyyy-MM-dd",
"yyyy-MM-ddTHH:mmK",
"yyyy-MM-ddTHH:mm:ssK",
"yyyy-MM-ddTHH:mm:ss.FFFFFFFK",
};
byte[] bytes = await new HttpClient().GetByteArrayAsync("https://example.com/sitemap.xml");
int urls = 0;
using var stream = new MemoryStream(bytes);
using var reader = XmlReader.Create(stream, settings);
while (reader.Read())
{
if (reader.NodeType != XmlNodeType.Element || reader.NamespaceURI != Ns) continue;
if (reader.LocalName == "loc")
{
urls++;
}
else if (reader.LocalName == "lastmod")
{
var v = reader.ReadElementContentAsString().Trim();
// TryParse would accept 03/14/2026. TryParseExact against the four
// permitted layouts is what the protocol actually asks for.
if (!DateTimeOffset.TryParseExact(v, w3c, CultureInfo.InvariantCulture,
DateTimeStyles.None, out _))
{
Console.Error.WriteLine("Invalid lastmod: " + v);
}
}
}
Console.WriteLine(urls + " URLs, " + bytes.Length + " bytes uncompressed");
if (urls > 50_000) Console.Error.WriteLine("Over the 50,000 URL limit");
if (bytes.Length > 52_428_800) Console.Error.WriteLine("Over the 50 MB limit");# The protocol publishes an XSD, so xmllint can check the structure offline.
curl -sL https://example.com/sitemap.xml -o sitemap.xml
curl -sO https://www.sitemaps.org/schemas/sitemap/0.9/sitemap.xsd
xmllint --noout --nonet --schema sitemap.xsd sitemap.xml
# The two hard limits:
xmllint --nonet --xpath 'count(//*[local-name()="loc"])' sitemap.xml # cap 50,000
wc -c < sitemap.xml # cap 52,428,800
# A gzipped sitemap must be decompressed before either check:
curl -sL https://example.com/sitemap.xml.gz | gunzip > sitemap.xml
# Pull every URL out, one per line, for a crawler or a spreadsheet:
xmllint --nonet --xpath '//*[local-name()="loc"]/text()' sitemap.xml | tr -s '\n'The security flags are not decoration: Java, Python and C# all resolve external entities unless told otherwise, and a sitemap is usually a file you did not write, arriving over a network from a plugin. The xmllint --nonet flag does the same job.
Common questions
Search Console says my sitemap has errors, but this tool says it is fine. Why?
The two look at different things. This page checks the file: structure, namespace, limits, dates, and the host and directory rules once you supply the sitemap address. Search Console also checks what only a crawler can, starting with whether the file was reachable.
"Couldn't fetch" usually means a 404, a redirect chain, a robots.txt rule covering the sitemap path, or an HTML error page served with a 200 status. Fetching the URL here shows the real status code. The other common mismatch is that the URLs are valid but the pages behind them are noindex or gone, which is a crawl problem.
Does Google use changefreq and priority?
No. Google states in its sitemap documentation that it ignores both. This tool reports their presence as information rather than an error, because they are legal elements and they do no harm.
Practically: no arrangement of priority values will make one page crawled ahead of another, and changefreq set to hourly will not raise your crawl rate. If you maintain code that calculates them, that code is not earning its keep. lastmod is the exception, and it is the one optional element worth getting right.
What format does lastmod need to be in?
W3C Datetime, a subset of ISO 8601. The simplest valid form is a plain date, 2026-03-14. If you include a time you must include a zone: 2026-03-14T09:30:00Z and 2026-03-14T09:30:00+00:00 are both correct.
The rejected forms are the ones databases and templates produce naturally. A space instead of the T is invalid, a timestamp with no zone is invalid, and so is a Unix epoch number or any day-first date such as 14/03/2026. If the bad dates come from a plugin rather than your own code, the value quoted in the error usually identifies which one.
How many URLs can one sitemap have?
50,000, and it must also stay under 52,428,800 bytes uncompressed. Whichever you hit first is the one that matters, and on a site with long URLs the size limit usually arrives before the count.
Past either, split the file and list the parts in a sitemap index: a small file in the same namespace whose children are <sitemap> elements, each with one <loc>. An index may list 50,000 sitemaps and is subject to the same 50 MB. Submit the index rather than each child, and reference it from robots.txt.
Why did the Fetch button fail on my sitemap?
Almost always CORS. Your browser will not let this page read a response from another domain unless that domain sends an Access-Control-Allow-Origin header, and few servers send one for sitemap.xml. That is your server's policy, not a fault in the file.
We do not work around it with a proxy, since a proxy would carry your URL and your file through a server we run. The failure message gives you a curl command for the address you entered. A gzipped sitemap needs decompressing first: curl -sL https://example.com/sitemap.xml.gz | gunzip.
Does my sitemap get uploaded anywhere?
No. The parser and every rule are JavaScript running in this tab. There is no back end to send anything to and no analytics with access to the editor. Open your network panel and validate a file: the page assets load once and then nothing.
A published sitemap is public, but the file you are debugging often is not. Staging sitemaps list unlaunched pages, internal tools and client sites under NDA, and several of the checkers ranking for this search are server-side. Your input is kept in this browser's localStorage so a refresh does not lose it; Clear removes it.