How this was measured
On 2026-09-29 we requested https://<domain>/sitemap.xml for the same 78 hand-picked domains used by our other surveys — news, commerce, SaaS, developer tooling, social, streaming, reference, SEO, education, finance, cloud, local and government — then read each domain's robots.txt and followed any Sitemap: line it declared. Every request was attempted twice, once directly and once through a local proxy, because a number of these domains are unreachable from our network without one.
A file counted as a sitemap only if it returned 200 and the body was XML rather than HTML or empty. Five of them arrive as gzip files (.xml.gz), which curl --compressed does not decompress — that flag handles transport encoding, not a file that is itself a .gz. Those five had to be gunzipped before they could be read at all, and our first pass stored them as binary garbage.
Then the part that is not usually measured. From each of the 55 sitemaps we extracted five URLs at even intervals and requested them: 251 requests in total. Every 3xx was re-requested with a real GET before being counted, because the first pass used HEAD to save bandwidth and HEAD does not always answer the same way GET does.
Finding 1: 55 of 78 serve one, and 25 of those are not where you would look
55 of 78 domains (70.5%) serve a readable sitemap. But only 30 of those 55 answer at /sitemap.xml. The other 25 are reachable only by following the Sitemap: line in robots.txt. If you audit sitemap coverage by requesting /sitemap.xml, you undercount by 45%.
The declared locations are not variations on a theme. trello.com points at https://a594014.sitemaphosting7.com/4704043/sitemap_4704043.xml — a third-party sitemap hosting domain, not a Trello URL at all. netflix.com uses /sitemap/index, with no file extension. wise.com uses /sitemap. wikipedia.org points at a MediaWiki REST endpoint, /w/rest.php/site/v1/sitemap/0, which answered 403 to us. bbc.com declares https://www.bbc.com/afrique/sitemap.xml, its French edition. ebay.com declares /lst/AUCTION-0-index.xml.
The 23 that did not produce a file break down as follows. 12 return a clean 404: etsy.com, linkedin.com, stackoverflow.com, python.org, rust-lang.org, go.dev, mozilla.org, w3.org, wikimedia.org, mit.edu, yelp.com and wikipedia.org. 4 return 200 with an HTML page: reddit.com (an 8 KB app shell), pinterest.com (1.36 MB of it), khanacademy.org (a 3 KB bot challenge) and cdc.gov, which redirects /sitemap.xml to /index.html and serves the homepage with a 200. 1 returns 200 with an empty body: x.com. 5 answer with an error: amazon.com 500, github.com 406, quora.com 403, npmjs.com 403, searchengineland.com 403. One, washingtonpost.com, was unreachable on both paths.
Finding 2: what robots.txt promises is not always what is there
74 of 78 domains serve a readable robots.txt; 55 of them declare a sitemap; 50 of those declarations resolve. Five do not:
washingtonpost.com declares a .xml.gz we could not reach. x.com declares /sitemap.xml, which is the empty 200 above. pinterest.com declares a seasonal-events file — seasonal_events_sitemap_20261129_first_sunday_of_advent_www.pinterest.com.xml — that returns 404, a generated reference left behind after the file it pointed at was retired. wikipedia.org declares the REST endpoint, 403 to us. searchengineland.com declares /sitemap_index.xml, also 403.
The failure runs the other way too. 5 sites serve a working sitemap and never declare it: shopify.com, gitlab.com, ahrefs.com, screamingfrog.co.uk and nih.gov. A crawler that only reads robots.txt misses those entirely.
Finding 3: most large sites ship an index, and 5 of them ship it gzipped
38 of the 55 are sitemapindex files pointing at children; 17 are flat urlset. The median index references 23 child files and is 2,584 bytes. The median flat urlset carries 617 URLs in 449 KB. The largest single file is gitlab.com at 19,009 URLs in 2.95 MB — inside the 50,000 URL / 50 MB limit, but one more product section away from needing to split.
Five serve their sitemap as a gzip file: nytimes.com, target.com, twitch.tv, vimeo.com and airbnb.com. Two omit the <?xml declaration entirely: theguardian.com and wise.com. Both still parse; neither needed to take the risk.
Four point at hosts other than their own. paypal.com declares 1,214 child sitemaps, all on www.paypalobjects.com, its CDN. notion.so lists 169 children on www.notion.com. nodejs.org includes three URLs on openjsf.org domains. youtube.com includes three on about.youtube and sibling hosts. who.int is the outlier: it declares its index over http:// and references children like http://www.who.int/SiteMaps/sitemap_static.xml.gz, while the site itself serves HTTPS.
Finding 4: lastmod is the field sites get wrong
18 of 55 sitemaps (33%) carry no lastmod at all. Three carry it inconsistently: asana.com on 611 of 3,487 entries, squareup.com on 5 of 16, and nih.gov on 2,688 of 2,689 — one entry short of complete.
Where the field exists, it is mostly date-only: 34,952 date-only values against 18,593 with a time. Date-only is valid and is what most of this sample chose; the ones carrying microseconds (cnn.com at 2026-09-28T23:24:32.775000+00:00) are the exception.
One file carries a value that is not a date. searchenginejournal.com lists -0001-11-30T00:00:00+00:00 alongside real entries like 2026-09-28T20:48:12+00:00. That is a zero-value timestamp leaking out of a templating layer — we cannot say which stack produced it, but a year of -0001 is not a date anyone set. It has the right shape, so a length check passes it, and it tells a crawler the page was last modified two thousand years ago.
Meanwhile 8 of 55 still emit changefreq and 6 still emit priority — fields Google has said publicly it ignores. The volumes are not small: nodejs.org writes 1,651 changefreq values, nih.gov writes 2,689 of them and 1,894 priority values, cloudflare.com writes 910 of each, shopify.com 557, theguardian.com 423, trello.com 112, stanford.edu 27, bbc.com 5.
Finding 5: one in eight sampled URLs does not answer 200
This is the number that matters, and it is the one nobody publishes. Of 251 URLs sampled from 55 sitemaps:
178 (70.9%) returned 200. 27 (10.8%) returned a redirect, confirmed by re-requesting with GET. 3 (1.2%) returned 404. 43 (17.1%) returned 401, 403 or 406 — and those are not sitemap faults. Six domains blocked every URL we tried: nytimes.com, reuters.com, bloomberg.com, ebay.com, coinbase.com and uber.com. Paywalls and bot walls, excluded rather than counted against them.
Excluding those six leaves 221 URLs: 178 (80.5%) returned 200, 27 (12.2%) redirected and 3 (1.4%) were dead. Put another way: 35 of the 55 sitemaps returned 200 for all five samples, and 12 of the 55 had at least one URL that redirects or 404s.
What the redirects are is more useful than how many there are:
walmart.com: 5 of 5. Every sampled /cp/ URL 301s to a /browse/ URL — /cp/blue-party-supplies/4514453 → /browse/party-occasions/blue-party-occasions/2637_1042319_1895000_4514453. That is a category URL migration the sitemap never absorbed.
stripe.com: 5 of 5. /legal → https://stripe.com/cn/legal. A geo redirect: we fetched from a Chinese IP address. The URL in the sitemap is not the URL a crawler in another country receives, and not the one that gets indexed.
zoom.us: 4 of 5. explore.zoom.us/es/customer_stories/eunis2020/ → www.zoom.com/es/customer-stories/all/. A domain migration, and one path changed shape on the way (/products/zoom-phone/ → /products/voip-phone/).
semrush.com: 4 of 5. /api-analytics/ → developer.semrush.com/api/v3/analytics/basic-docs/ — off the main domain entirely.
paypal.com: 2 of 5. /ae/webapps/mpp/mobile-apps?locale.x=ar_AE → /ae/digital-wallet?locale.x=ar_AE. The path was replaced; the query string survived.
cloudflare.com and stanford.edu: trailing slashes, in opposite directions. /case-studies/fossil 301s to /case-studies/fossil/; /health-medicine/ 308s to /health-medicine. Neither site agrees with the other about which form is canonical, and the sitemap disagrees with both.
The three genuine 404s: gitlab.com/gitlab-org/security-products/dependencies/golang/toml/-/issues, and two legacy release posts at nodejs.org/en/blog/release/v08.20 and /v2013.0.
Redirects outnumber dead URLs nine to one. That ratio matters because the two failures need different fixes. A 404 is a deletion — remove the entry. A redirect is a stale reference you have to trace to its target and then decide whether the sitemap or the URL is the thing that should change.
What to do with this
Audit the URLs, not the file. A sitemap that parses cleanly tells you nothing about whether its contents resolve. Sample ten URLs from yours and request them. Twelve of the 55 sitemaps here would have failed that test, and every one of them is a sitemap that "works".
Do not assume /sitemap.xml. Read the Sitemap: line in your robots.txt and request that URL. Twenty-five of the 55 sitemaps in this sample live somewhere else, and five of them are on paths with no .xml extension. Our robots.txt tester reads the file and shows you which directives a crawler would apply.
Regenerate after every URL migration. Walmart's /cp/ entries are what an un-regenerated sitemap looks like at scale: every sampled URL pointing at a structure that no longer exists. The redirect saves the user; it does not save the crawl budget.
If you geo-redirect, test from more than one country. Stripe's /cn/legal is a correct redirect doing its job, and it means the URL in the sitemap is not the URL that gets indexed from most of the world.
Pick a trailing-slash convention and make the sitemap match it. Cloudflare and Stanford chose opposite conventions, and both sitemaps disagree with their own site.
Either put lastmod on every entry or on none. A third of this sample ships none and is not visibly worse off. Three ship partial coverage, which is the version with no upside.
Drop changefreq and priority. Eight sites here write thousands of values Google ignores. If you serve .xml.gz, remember that the file is gzip, not just the transfer.
Limitations
78 hand-picked well-known domains, not a random sample of the web. Categories are our own manual grouping, several contain five domains, and one site moves a rate by 20 points.
All requests came from one IP address in China on one day, and that is directly visible in the data: 43 of 251 sampled URLs answered 401, 403 or 406, six domains blocked everything we tried, and Stripe redirected us to its /cn/ edition. Those are excluded rather than counted as defects. A different vantage point would find more 200s and fewer redirects.
Five URLs per sitemap, evenly spaced. A sitemap with one bad URL among 19,009 — GitLab's case — has roughly a 0.03% chance of showing it in our sample. These numbers describe the sample, not the full files, and the true per-sitemap defect rate is certainly higher than 12 of 55.
17 of the 38 child sitemaps we fetched were truncated at 400 KB before parsing, so their entry counts are lower bounds.
The 3xx count was confirmed with GET, but the initial pass used HEAD wherever it worked. Six URLs that looked like redirects turned out to be 403s on GET, and none turned out to be 200s — so the confirmed figure is lower than the raw one. If another host answers HEAD differently, this method over-counts redirects.
We did not follow redirect chains and did not check whether the redirect target is itself canonical, so "redirects" here includes both harmless normalisation and real migrations.
We read files discoverable from robots.txt or /sitemap.xml only. A sitemap submitted through Search Console and declared nowhere is invisible to us — and the five sites serving a sitemap with no declaration are proof that this configuration exists.
Reproduce it
The script and the shared domain list ship with this site. node scripts/survey-sitemap.mjs fetches and stores every response; --children pulls the first child of each index; --probe samples five URLs per sitemap; --recheck re-requests every 3xx with GET; --report re-derives every number from stored data without touching the network; --evidence prints the raw XML behind each verdict. Our robots.txt, canonical, hreflang, AI crawler and llms.txt surveys use the same domain list from different angles.