XML sitemaps that actually work

The sitemaps protocol is one web page long and defines six elements. That should make it hard to get wrong, and yet the files search engines quietly drop fail on the same short list: a namespace that is nearly right, an ampersand nobody escaped, a lastmod stamped with the build time, and a file that crossed 50,000 URLs three deploys ago.

What follows is the protocol as written, the two limits that are hard failures, the elements Google states in its own documentation that it ignores, and the one it uses only when you have earned its trust. There is a numbered checklist at the end.

#The protocol, and the namespace that has to be exact

Sitemaps 0.9 was published at sitemaps.org in 2006 and has not changed since. It is a list of URLs with optional metadata, and the protocol is explicit about its own standing: using it "does not guarantee that web pages are included in search engines". It excludes nothing either, so a URL absent from your sitemap can still be crawled and indexed.

The namespace URI has to be written exactly as http://www.sitemaps.org/schemas/sitemap/0.9. That is http rather than https, and there is no trailing slash. A namespace URI is an identifier compared character by character, not an address anything fetches, so the https variant produces a perfectly well-formed document in a vocabulary no search engine recognises. Nothing about the file looks wrong, which is why this one survives review.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/catalog/item?id=12&amp;sort=asc</loc>
    <lastmod>2026-08-14T09:30:00+00:00</lastmod>
  </url>
  <url>
    <loc>https://example.com/about/</loc>
    <lastmod>2025-11-02</lastmod>
  </url>
</urlset>
A complete sitemap. Everything below <loc> is optional.
ElementCardinalityRule
urlsetRequired, rootMust carry the namespace declaration.
urlRequired, 1 to 50,000One per URL. A parent only; it holds no text.
locExactly one per urlAbsolute, including the scheme. Under 2,048 characters.
lastmod0 or 1W3C Datetime. A bare YYYY-MM-DD is legal.
changefreq0 or 1One of seven lowercase values. Ignored by Google.
priority0 or 10.0 to 1.0, default 0.5. Ignored by Google.
Every element the protocol defines, with its cardinality.

The seven changefreq values are always, hourly, daily, weekly, monthly, yearly and never, all lowercase. Image, video, news and hreflang data are not in this list; they arrive through separate namespaces, covered below.

#Two hard limits, and the index file that gets you past them

A single sitemap may contain no more than 50,000 URLs and must be no larger than 50 MB, which the protocol pins to the exact figure 52,428,800 bytes. Both are ceilings, not guidance. Cross either and the file is rejected in full rather than truncated, so a site that grows past the line loses every URL in it at once.

Past either limit you split the file and list the parts in a sitemap index: same namespace, root element sitemapindex, children that are sitemap elements rather than url elements. Mixing the two child types in one document is invalid. An index is itself capped at 50,000 entries and 50 MB, which puts the ceiling for one index at 2.5 billion URLs.

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://example.com/sitemaps/products-1.xml.gz</loc>
    <lastmod>2026-08-14</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemaps/products-2.xml.gz</loc>
    <lastmod>2026-06-02</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemaps/articles.xml.gz</loc>
    <lastmod>2026-08-14</lastmod>
  </sitemap>
</sitemapindex>
sitemap-index.xml, served at the site root

The lastmod on a sitemap entry inside an index is the most useful date in the system and the most commonly wasted. It tells a crawler which child files are worth re-fetching. If products-2.xml has not changed since June, saying so lets the crawler skip 50,000 URLs it has already seen; if your build stamps all three with today, it has to open everything.

Split by something stable, such as content type or publication month, rather than by an arbitrary slice of a database cursor, because chunk boundaries shift on every build and make every file look changed. And do not nest an index inside an index: Google does not document support for it.

#lastmod, changefreq and priority

lastmod takes a W3C Datetime, the ISO 8601 profile described in the W3C note "Date and Time Formats". Valid forms run from a bare date, 2026-08-14, to a full timestamp with a timezone, 2026-08-14T09:30:00+00:00. A timestamp with no zone is not valid.

Google states that it uses lastmod "if it's consistently and verifiably accurate", and is specific about what should move the date: a significant update to the main content, the structured data or the links. Fixing a typo does not count. Rolling the copyright year does not count. "Verifiably" is doing real work in that sentence, because Google can compare your claim against what it fetched last time.

Every URL stamped with the build time
<url>
  <loc>https://example.com/about/</loc>
  <lastmod>2026-08-14T03:14:07+00:00</lastmod>
</url>
<url>
  <loc>https://example.com/pricing/</loc>
  <lastmod>2026-08-14T03:14:07+00:00</lastmod>
</url>
The date the content actually changed
<url>
  <loc>https://example.com/about/</loc>
  <lastmod>2024-03-19</lastmod>
</url>
<url>
  <loc>https://example.com/pricing/</loc>
  <lastmod>2026-08-11</lastmod>
</url>

The left-hand version is what a static site generator produces when nobody thinks about it, and it is worse than omitting lastmod entirely: it asserts that every page on the site changed at 03:14 this morning. The first time a crawler re-fetches half a dozen and finds identical bytes, the signal is spent for the whole file rather than for the URLs that lied. If your generator cannot reach a real modification date, git log -1 --format=%cI on the source file is usually the honest answer; leaving the element out is the next best.

changefreq and priority need less discussion, because Google's "Build and submit a sitemap" documentation says flatly that it ignores both. changefreq was always described in the protocol itself as "a hint and not a command". priority defaults to 0.5 and is explicitly relative to other URLs on your own site, so setting every page to 1.0 encodes exactly what setting every page to 0.5 encodes. Neither element is an error, but a CMS computing a priority score for 40,000 URLs is doing arithmetic that nothing reads.

#Where a sitemap is allowed to live

A sitemap can only vouch for URLs at or below its own directory. The protocol gives the example directly: a sitemap at http://example.com/catalog/sitemap.xml may include any URL starting with http://example.com/catalog/ but may not include URLs starting with http://example.com/images/. Scheme, host and port must match too, so a sitemap served over https cannot list http URLs, and www.example.com and example.com are different hosts here.

That is the argument for putting the file at the site root and stopping there. One at /assets/seo/sitemap.xml covers /assets/seo/ and nothing else, and the URLs outside that prefix are dropped silently rather than reported.

Google adds a second route: sitemaps submitted through Search Console are not bound by the directory rule, and an index may reference files on another host once cross-site submission is configured there. Prefer the root of the host the sitemap describes to either.

The sitemap validator here asks for the URL you intend to publish at, and enables the host, scheme and directory-prefix checks when you supply one. Without it, those rules are listed under "not checked" rather than passed, because a rule that could not run must never render as a rule that succeeded. The same principle sorts the rest: a priority outside 0.0 to 1.0 is an error, while merely using priority at all is reported as information with a count.

#Escaping, in the right order

A raw ampersand in a query string is the most common thing wrong with sitemaps in the wild, and it is worth being precise about what it breaks. It is not a sitemap error. It is a fatal XML well-formedness error, so the parser stops and the entire file is discarded, not the one URL that contained it. A sitemap with 30,000 good URLs and one unescaped ampersand indexes nothing.

Fatal. The parser stops here.
<loc>https://example.com/s?q=nuts & bolts&page=2</loc>
Percent-encoded, then XML-escaped
<loc>https://example.com/s?q=nuts%20%26%20bolts&amp;page=2</loc>

The order of operations is what people get wrong. Percent-encode first, treating the URL as a URL under RFC 3986: a space becomes %20, and an ampersand that is part of a value rather than a separator becomes %26. Then XML-escape the result, which only ever affects the five characters below. The other order produces %26amp%3B in your URLs, and doing both to the same character produces &amp;amp;, which crawlers fetch literally.

CharacterWritten as
&&amp;
'&apos;
"&quot;
>&gt;
<&lt;
The five characters that must be escaped inside <loc>.

The file must be UTF-8 and the declaration must say so. Non-ASCII in a URL is handled before the XML layer: an internationalised host name goes to punycode, and non-ASCII in the path is percent-encoded from its UTF-8 bytes. Python's urllib.parse.quote and JavaScript's encodeURIComponent both do the path half correctly, and neither escapes XML afterwards, which is the step generators skip.

#The image, video and news extensions

Three Google extensions add elements inside url, each in its own namespace, each declared on the root element. Using a prefix that has not been declared is a namespace error and fails the whole file, exactly like the ampersand.

ExtensionNamespace URIKey constraints
Imagehttp://www.google.com/schemas/sitemap-image/1.1image:image wraps image:loc. Up to 1,000 images per url. image:caption, image:title, image:geo_location and image:license were deprecated in spring 2022 and are no longer read.
Videohttp://www.google.com/schemas/sitemap-video/1.1Requires video:thumbnail_loc, video:title, video:description (2,048 characters maximum) and at least one of video:content_loc or video:player_loc. video:duration is 1 to 28800 seconds, video:rating is 0.0 to 5.0, and at most 32 video:tag per video.
Newshttp://www.google.com/schemas/sitemap-news/0.9news:news requires news:publication (holding news:name and news:language), news:publication_date and news:title. At most 1,000 news:news entries, and only articles from the last two days.
hreflanghttp://www.w3.org/1999/xhtmlxhtml:link rel="alternate" hreflang="en-gb" inside url. Every language variant must list every other variant, including itself.
Extension namespaces and the constraints that matter.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:image="http://www.google.com/schemas/sitemap-image/1.1"
        xmlns:video="http://www.google.com/schemas/sitemap-video/1.1">
  <url>
    <loc>https://example.com/reviews/kettle</loc>
    <lastmod>2026-07-30</lastmod>
    <image:image>
      <image:loc>https://cdn.example.com/kettle-1.jpg</image:loc>
    </image:image>
    <video:video>
      <video:thumbnail_loc>https://cdn.example.com/kettle.jpg</video:thumbnail_loc>
      <video:title>Kettle review</video:title>
      <video:description>Four minutes on descaling.</video:description>
      <video:player_loc>https://example.com/player?v=kettle</video:player_loc>
      <video:duration>247</video:duration>
    </video:video>
  </url>
</urlset>
Declaring the prefixes you use, and only those.

A news sitemap is the one with an expiry attached: only articles from the last two days belong in it, so it has to be regenerated on a schedule rather than at deploy time, with old entries removed rather than accumulated.

#The checklist

  1. The file parses as XML. Check this first, because every other check depends on it.
  2. The root is urlset (or sitemapindex) and carries xmlns="http://www.sitemaps.org/schemas/sitemap/0.9", with http and no trailing slash.
  3. The file is UTF-8, the declaration says UTF-8, and there is no byte order mark before it.
  4. Under 50,000 url entries and under 52,428,800 uncompressed bytes, with headroom. Split at 45,000 rather than at 49,999.
  5. Every url has exactly one loc: absolute, with a scheme, under 2,048 characters.
  6. Every loc is percent-encoded first and XML-escaped second, and no ampersand appears raw anywhere in the file.
  7. Every loc shares the sitemap's scheme, host and port and sits under its directory, or robots.txt cross-submission is set up.
  8. No duplicate loc values, and no URLs that redirect, 404, or are blocked by robots.txt or a noindex tag.
  9. lastmod, if present, is a valid W3C Datetime, is not in the future, and reflects a real content change rather than the build time.
  10. changefreq and priority are either correct or absent. Google reads neither, so absent is cheaper.
  11. Extension prefixes are declared on the root, and video and news entries carry their required children.
  12. The file is named in robots.txt with a Sitemap: line and submitted once in Search Console. Re-submitting on every deploy achieves nothing.

Common questions

Do I actually need a sitemap?

If your site is small, internally well linked, and every page is reachable in a few clicks from the home page, a sitemap adds very little. Google says as much in its own documentation: crawlers find pages by following links, and a sitemap is not a substitute for a navigable site.

It earns its place when the site is large, when pages are isolated or newly published with nothing linking to them yet, when the site is new and has few external links, or when you have rich media or news content that the extensions can describe.

Should I remove priority and changefreq from my existing sitemaps?

There is no urgency. Neither element causes harm, and other crawlers may still look at them, though none documents doing so.

The reason to remove them is size and honesty. Two extra elements per url is roughly 60 bytes, which is 3 MB across 50,000 URLs and a meaningful slice of the 50 MB budget. The stronger reason is that a priority score computed by your CMS creates the impression internally that someone is tuning something, when nothing downstream reads the number.

Should I gzip my sitemap?

Yes, for anything large. The protocol permits gzip, search engines accept a .xml.gz file, and XML compresses extremely well because it is so repetitive: a sitemap typically drops below a tenth of its size.

What gzip does not do is give you more room, since both limits are measured against the uncompressed file. Serve the compressed file with the gzip content type, or let your server negotiate the encoding; do not gzip the bytes and then serve them as text/xml.

Search Console says my sitemap "could not be read". Where do I start?

In rough order of how often it turns out to be each: the file is not well-formed XML, usually an unescaped ampersand; the URL you submitted returns a redirect or a 404 rather than the file; the response carries the wrong content type; or the namespace is subtly wrong, most often the https variant of the sitemaps.org URI.

Check the first locally before anything else, because it costs nothing and it is the answer most of the time. Then fetch the exact URL you submitted with curl rather than a browser, so you see the status code and the headers instead of a rendered result.

Sources

Try it