Technical

Sitemap

Sitemap is an XML or plain-text file listing the URLs of a site you want search engines to know about, optionally with last-modified dates. It is a hint about what exists, not an order to index.

How it is measured

Fetch `/sitemap.xml` and check for 200 and valid XML, with at most 50,000 URLs and 50 MB uncompressed per file. Bigger sites use a sitemap index that points at several files. Each `<loc>` should be a canonical, indexable URL returning 200.

Submit it in Search Console, which reports discovered and indexed counts for it. The gap between listed and indexed is the useful number.

Worked example

A WordPress store's sitemap lists 9,400 URLs and Search Console shows 6,100 indexed. Crawling the sitemap finds 1,900 URLs that redirect, 800 with noindex, and 600 that return 404. After the plugin is told to drop tag archives and out-of-stock redirects, the file shrinks to 6,100 URLs and indexed pages rise to 5,800.

Nothing was added to the site. The sitemap simply stopped listing pages that could never be indexed.

How it differs

A sitemap lists URLs to crawl. robots.txt lists paths to avoid. One invites and the other asks crawlers to stay out. A URL that appears in both is a contradiction that tools will flag.

Common errors

Listing redirected, noindexed, or 404 URLs. Listing non-canonical variants. Using the wrong host or an http scheme. Setting `lastmod` to the current time on every build. Exceeding the size limits. Forgetting to reference the file in robots.txt.

In practice

Crawl your own sitemap and fix every non-200 row. List only canonical, indexable URLs. Reference the file in robots.txt and submit it once in Search Console.

See also

robots.txt, Canonical URL

Sources

Count this on a real site.

Watch my website