XML sitemaps: what belongs in one and what doesn't
A sitemap is a hint, not an instruction. Which URLs belong, why listing a noindex page contradicts itself, and when lastmod is simply ignored.
· 5 min read
A sitemap is a hint, not an instruction
An XML sitemap is a file listing URLs on your site that you would like a search engine to know about. That is the whole of its function, and almost every misunderstanding comes from expecting more. It does not cause pages to be indexed. It does not make them rank. It does not compel a crawler to visit anything. It is a suggestion — a list saying these URLs exist and here is what I can tell you about them.
Which means a sitemap cannot fix the problems people most often deploy it against. A page that is not being indexed because it is thin will not be indexed because you listed it. A page nothing links to remains unsupported after being listed, because a sitemap entry asserts existence where a link asserts relevance. Sitemaps genuinely help in narrower circumstances: a large site where crawling might not reach everything by following links, a new site with few external links, a page deep in the structure, and content that changes often enough that you want the change noticed. For a twenty-page business site that is well linked internally, the sitemap is a small convenience rather than a lever.
What belongs in it
The rule is short: every URL you want indexed, in its canonical form, and nothing else. Canonical form is the part that gets skipped. If your site serves content at both the www and non-www hostname, list only the one you have settled on. If URLs work with and without a trailing slash, list the form your canonical tags name. If a page has a preferred URL and several parameter variants, list the preferred one. A sitemap listing a non-canonical variant tells an engine you consider that variant the URL worth knowing about, while your canonical tag says otherwise — and the contradiction is entirely self-inflicted.
So: your real pages, your articles, your products, your category pages where they are worth indexing. In canonical form, returning a successful response, and each appearing once. This sounds obvious and is routinely broken by generators that emit every URL the platform can produce rather than every URL you want indexed, which is a different set and usually a much larger one.
What does not belong, and why each one contradicts something
A page marked noindex should not be in the sitemap. Listing it says 'please know about this URL' while the page itself says 'do not list me', and giving an engine two contradictory instructions about one URL means one of them is being disregarded. Decide which you mean. Same for a URL blocked in robots.txt: you have asked a crawler not to fetch it and simultaneously invited it to.
A URL that redirects does not belong — list the destination instead, since a redirect entry wastes the crawl and asserts that a URL you have deliberately retired is one worth knowing about. A page canonicalised to another URL does not belong; list the canonical. A URL returning an error does not belong, and these accumulate on their own as pages are deleted without the sitemap being regenerated, which is why a stale sitemap full of dead URLs is a common finding. Pages requiring login do not belong, since a crawler will only ever receive the login screen. And URLs with tracking parameters do not belong, because each is a distinct address serving content already listed under its clean form. The unifying principle: a sitemap should agree with everything else your site says, and each of these entries makes it disagree with something.
lastmod, and when it is ignored
The lastmod element states when the URL's content last changed meaningfully, and it is the one optional element worth getting right, because Google has said it uses lastmod when it is accurate and disregards it when it is not. That conditional is the whole story. A site whose sitemap sets lastmod to today for every URL on every regeneration has produced a field carrying no information — if everything changed today, nothing did — and the sensible response is to stop trusting it, which is what happens.
So either populate lastmod from the actual date the page's substantive content changed, or leave it out. Leaving it out is a perfectly respectable choice, and a great deal better than an automated value that is technically present and semantically empty. Note also what counts as a change: a modification to the content a visitor reads, not a template tweak, a footer year update or a redeploy that touched every file. The two sibling elements, changefreq and priority, are worth mentioning only to say that Google ignores them. They were part of the original specification, they carry no weight, and time spent tuning priority values across a sitemap is time spent on a field nobody reads.
The limits nobody hits, and the ones people do
A single sitemap file may contain at most 50,000 URLs and must not exceed 50MB uncompressed. These are the numbers everyone quotes and almost no small business approaches — a fifty-page site is three orders of magnitude below the ceiling. If you do exceed either, the mechanism is a sitemap index file: a sitemap of sitemaps, listing several files each within the limits, which is also a tidy way to organise large sites by section so a problem can be localised.
The limits that actually cause trouble are different and unmeasured. A sitemap that is stale, listing pages deleted months ago and omitting pages published last week, because generation is manual and nobody remembers. A sitemap that was submitted once and has never been looked at since, quietly reporting errors. A sitemap listing every faceted URL a platform can generate, which buries the fifty pages you care about among thousands you do not. And a sitemap that is not discoverable, because it was never referenced in robots.txt and never submitted anywhere. None of these trips a limit. All of them are more common than a site with 50,000 pages.
Checking your own, for free
Nearly everything here is verifiable at no cost. Open your sitemap in a browser — conventionally at /sitemap.xml, and referenced by a Sitemap line in robots.txt, which is worth confirming is present. Read it. Then take a sample of the URLs it lists and actually request them, checking that each returns successfully rather than redirecting or erroring, and that each is the canonical form rather than a variant. Check a handful of pages against the noindex question: search the page's source for noindex and confirm no page listed in the sitemap carries it. Check whether lastmod values differ between URLs or are all identical, which tells you immediately whether the field means anything.
What needs a credential is the engine's own view. Search Console reports, for a verified property, when your sitemap was last read, how many of its URLs are indexed, and which entries produced errors — free, and unavailable any other way. That last figure is the one worth watching, because a large gap between submitted and indexed is not a sitemap problem to be fixed by resubmitting; it is a signal that the pages themselves are not being judged worth a slot. Resubmitting a sitemap in that situation is the most common wasted action in this whole subject, because the sitemap was never the constraint.
Common questions
Will adding a page to my sitemap get it indexed?
No. A sitemap makes a URL known, which is different from making it worth indexing. If a page is being passed over because it is thin or duplicative of your other pages, listing it changes nothing. Resubmitting a sitemap when the gap is between submitted and indexed counts is the most common wasted action here.
Should I include noindex pages so Google finds out about the noindex?
No. The sitemap says the URL is worth knowing about and the page says do not list it, which is a contradiction that gets one of the two disregarded. If you want a page dropped, keep it crawlable so the tag can be read, and leave it out of the sitemap.
Do changefreq and priority do anything?
Google ignores both. They came from the original sitemap specification and carry no weight, so tuning priority values across a large sitemap is effort spent on a field nobody reads. Lastmod is the one optional element that matters, and only when it reflects real content changes.
How many sitemaps do I need?
One, until you exceed 50,000 URLs or 50MB uncompressed, which almost no small business does. Beyond that, use a sitemap index file listing several sitemaps. Splitting by section is also a reasonable choice for a large site, because it makes an error easier to localise.
Related pages