Free tool
Free Sitemap Validator
Validate any XML sitemap or sitemap index, count every URL it lists, and check that robots.txt points search engines to it. Free, instant, no signup.
Quick answers
What is a sitemap validator?
A sitemap validator checks that your XML sitemap is well-formed, follows the sitemaps.org protocol, and can be fetched by search engines. It parses the file, follows every child sitemap in a sitemap index, counts the URLs, and flags the errors that stop Google from reading it.
Search Console only tells you something is wrong after Google has tried to read the file, sometimes days later. A sitemap checker gives you the answer in seconds, before you deploy or right after, so a broken sitemap never gets the chance to hide your new pages.
Why is a sitemap important for SEO?
Because Google can only index pages it knows exist. In our indexing study, comparing pages on the same websites, Google had never discovered 59.6% of the pages left out of the sitemap, against 17.1% of the pages listed in it.
The sitemap is the one place where you can hand search engines a complete list of the pages you want indexed. Pages outside it have to be found by following links, and on the same websites, most of them never were. The full numbers are further down this page.
How do I check if my sitemap is valid?
Enter your domain or sitemap URL in the validator above and click Validate Sitemap. It finds the sitemap through robots.txt or the usual paths, parses it and every nested sitemap, counts the URLs, and lists errors and warnings, including a sitemap missing from robots.txt.
To test a sitemap before it goes live, switch to the Paste XML tab and paste the raw file. Once it is live, the Sitemaps report in Google Search Console shows whether Google could read it and how many pages it discovered, but only after Google has fetched it.
How to use this sitemap validator
Testing a sitemap takes a few seconds and works on any public website, including a competitor's, or on a file you have not deployed yet. No account, no Search Console access needed.
Enter your domain or the full sitemap URL, or switch to Paste XML to check a file before you deploy it
Click Validate Sitemap. The tool finds your sitemap through robots.txt or common paths and follows every child sitemap
Fix the errors first, then the warnings, and run the check again until the sitemap comes back clean
Original research
Why your sitemap decides what Google indexes
Every SEO guide calls a sitemap good practice. We wanted a number. Using Google's URL Inspection API, we asked Google what it had done with 611,503 pages on 255 real websites, built on WordPress, Shopify, Webflow, Wix, Next.js and custom stacks, between May and September 2026. Of everything we tested, being in the sitemap had by far the biggest effect on whether a page got indexed.
- Websites studied
- 255
- Pages inspected by Google
- 611,503
- Indexation gain for pages in the sitemap
- +42 pts
- Pages Google never fetched
- 1 in 3
Pages missing from the sitemap mostly never get discovered
We compared pages on the same websites, so a strong site cannot flatter the result. On the 78 websites with at least 10 pages inside the sitemap and 10 outside it, the difference was anything but subtle:
| Same 78 websites | Never discovered by Google | Indexed once Google crawled it |
|---|---|---|
| In the sitemap | 17.1% | 88.1% |
| Not in the sitemap | 59.6% | 80.6% |
All page types, compared within each website.
That is a 42-point gap in getting found, against a 7.5-point gap once Google has actually crawled the page. A sitemap doesn't make Google like your pages more. It makes Google aware they exist. For blog articles, pages listed in the sitemap were indexed 42.4 percentage points more often than articles on the same site that were left out, and the gap held on 67 of the 78 websites.
For comparison, the classic on-page checklist barely registered in the same study. Word count, title length and Title Case made no measurable difference to whether a page got indexed.
Internal links alone are not enough
The common advice is that a well-linked site doesn't need a sitemap, because crawlers follow links. Google's own documentation says a small, comprehensively linked site might get by without one. Our data points the other way. A page that is not in the sitemap can only be discovered by following a link to it, and on the same websites, Google had never discovered 6 in 10 of those pages.
Internal links still matter. They tell Google which pages are important, help it crawl deeper, and pass authority to new content. But discovery is a queue, and a crawler only works through part of it on each visit. A sitemap puts every URL in the queue on the day you publish, instead of waiting for a crawl to stumble onto it. Use both: link to new pages from pages that are already indexed, and list every page you want indexed in the sitemap.
The more of your site is in the sitemap, the more of it gets indexed
Group websites by sitemap coverage, the share of their pages listed in their own sitemap, and indexation climbs step by step:
20 websites
27 websites
45 websites
74 websites
166 websites and 43,296 pages. Middle website in each group.
Fit a straight line through 246 websites and an empty sitemap predicts 23% of pages indexed, a complete one 67%. Sitemap coverage explained 16.3% of the difference in indexation between websites. Domain Rating explained 4.5%, and site size 0.6%. What you put in your sitemap predicted indexation more than three times better than your backlink authority did.
Incomplete sitemaps are common, too. 110 of the 362 websites we checked had more than 20% of their pages missing from their own sitemap, and 48 were missing more than half.
Most unindexed pages were never crawled at all
Follow the pages of the average website through Google's pipeline and the biggest loss happens before Google reads a single word:
Every page meant to be indexed
26.8% were never discovered
5.6% were discovered but never crawled
7.5% were crawled, then left out
Average of 246 websites with at least 20 pages, each website weighted equally.
Roughly one page in three had never been fetched by Google, not even once, while only 7.5% were crawled and then turned down. Google cannot reject a page it has never read. So before you rewrite a page that isn't indexed, check that it's in your sitemap and that Google has crawled it.
Your CMS changes discovery, not quality
Platforms differ a lot in how many pages Google never discovers, and hardly at all in what happens once Google reads a page:
| Platform | Websites | Never discovered | Indexed once crawled |
|---|---|---|---|
| Shopify | 24 | 16.0% | 88.1% |
| WordPress | 91 | 22.1% | 89.5% |
| Webflow | 9 | 26.6% | 88.0% |
| Next.js | 29 | 28.8% | 90.9% |
| Custom or unknown | 82 | 32.2% | 87.8% |
| Wix | 6 | 45.6% | 87.1% |
Websites with at least 20 pages, each website weighted equally. Webflow and Wix are small samples, so read those rows as a signal rather than a verdict.
The share of pages Google never discovered ranges from 16% to 46% depending on the platform, while the share indexed once crawled stays between 87% and 91% everywhere. Shopify and WordPress generate sitemaps out of the box. On Next.js and custom stacks, the sitemap is whatever your code outputs, so a forgotten route or a stale build quietly leaves pages out, and Google never hears about them.
Give new pages 30 days before you judge them
Timing matters as much as the sitemap. On 68 websites with at least 10 articles checked at both moments, the typical site had 8% of its articles indexed when Google was asked within 7 days of publishing, and 90% when asked 31 to 90 days after. 80% of the articles checked in their first week came back as "URL is unknown to Google". A week-one report measures Google's queue, not your content.
How we ran the study
We queried Google's URL Inspection API through Search Console for 611,503 pages on 255 websites that connected Search Console to BlogSEO, between 3 May and 8 September 2026. Every website counts once, so a store with 88,000 pages weighs the same as a blog with 40. The in-sitemap versus out-of-sitemap comparison uses the 78 websites with at least 10 pages on each side and compares pages within each website. Requiring only 3 pages per side (111 websites) gives +40.7 points, requiring 20 (53 websites) gives +45.5. This is observational data from websites that use an SEO product, not a controlled experiment: it shows a strong association, not proof of cause.
XML sitemap format, with examples
An XML sitemap is a UTF-8 encoded file, usually at /sitemap.xml, that lists the URLs you want search engines to crawl. A minimal valid sitemap looks like this:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/</loc>
<lastmod>2026-09-20</lastmod>
</url>
<url>
<loc>https://example.com/blog/xml-sitemap-guide</loc>
<lastmod>2026-09-24</lastmod>
</url>
</urlset>Each <url> entry needs a <loc>. The optional tags behave differently than most guides suggest:
<loc>is required: the full, absolute URL including https://, on the same host as the sitemap, with special characters escaped (an ampersand becomes&).<lastmod>is optional, and the only optional tag Google uses. It trusts the date when it is consistently accurate, so only change it when the content of the page actually changes.<changefreq>and<priority>are ignored by Google. You can leave them out.
One sitemap file can hold up to 50,000 URLs and 50 MB uncompressed. Past that, split the URLs into several files and list them in a sitemap index:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.com/sitemap-pages.xml</loc>
</sitemap>
<sitemap>
<loc>https://example.com/sitemap-articles.xml</loc>
<lastmod>2026-09-24</lastmod>
</sitemap>
</sitemapindex>Search engines read the index, then fetch each child sitemap. This XML sitemap validator does the same: it follows nested indexes up to three levels deep, reads up to 50 child sitemaps per index, and adds up the URL count across all of them. Gzip-compressed sitemaps (.xml.gz) work too.
What this sitemap checker tests
Every check runs the way a search engine would see your sitemap, as an anonymous visitor fetching public URLs:
- Sitemap discovery. When you enter a domain, the tool reads the Sitemap line in robots.txt, then tries /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml and /wp-sitemap.xml.
- Valid XML. The file must parse, with a
<urlset>or<sitemapindex>root element. - Sitemap indexes. Every child sitemap is fetched and parsed, and any that fail to load are reported.
- URL count. The total number of URLs across all sitemaps, with a sample so you can spot the wrong domain or protocol at a glance.
- Content type. A warning when the server returns something other than XML.
- Empty sitemaps. A sitemap that parses but lists zero URLs, a classic sign of a broken plugin or build step.
- robots.txt. Whether robots.txt is reachable, declares a sitemap at all, and declares this one.
Common sitemap errors and how to fix them
These are the problems that most often keep a sitemap from working, whether the validator flags them or Search Console does.
Sitemap could not be read
Why it happens: Search Console shows this, or "Couldn't fetch", when Googlebot cannot download or parse the file. The usual culprits are a 404, 403 or 5xx status, a redirect, a firewall or bot protection that blocks Googlebot, or an HTML page served where the XML should be. For a sitemap you submitted in the last few days, it can also simply mean Google has not fetched it yet.
How to fix it: Open the sitemap URL in a private window, then run it through the validator above. It should return a 200 status with an XML content type, and robots.txt should not block it.
Invalid XML
Why it happens: The file does not parse. Unescaped ampersands in URLs are the most common reason, followed by a stray character or blank line before the XML declaration, and PHP or server errors printed into the file.
How to fix it: Escape special characters in URLs (& becomes &), make sure the file starts with the <?xml declaration, and save it as UTF-8.
Not a sitemap
Why it happens: The URL returns valid markup, but the root element is not <urlset> or <sitemapindex>. It is usually an HTML page, a soft 404, or a redirect to the homepage.
How to fix it: Point robots.txt and Search Console to the real sitemap file. If you are not sure where it lives, enter your domain above and the validator will look for it.
Wrong content type
Why it happens: The server sends the sitemap as text/html or text/plain, often because a CMS route or a CDN rule treats it like a regular web page.
How to fix it: Serve sitemap files with an application/xml or text/xml content type.
Child sitemap fails to load
Why it happens: A sitemap index lists a child sitemap that returns an error, was renamed, or times out. Search engines cannot read any of the URLs listed in that child sitemap.
How to fix it: Regenerate the index so it only lists live files, then validate each failing child sitemap URL on its own.
Empty sitemap
Why it happens: The file parses but lists zero URLs, typically after a plugin update, a disabled content type, or a build step that ran before the content was available.
How to fix it: Check the sitemap settings in your CMS or SEO plugin, confirm which content types are included, and rebuild the file.
Sitemap missing from robots.txt
Why it happens: robots.txt has no Sitemap line, or it still points to an old sitemap URL, so crawlers that never saw your Search Console submission cannot find the file.
How to fix it: Add a line such as Sitemap: https://yoursite.com/sitemap.xml with the full, absolute URL. You can list several sitemaps, one per line.
URLs that should not be in the sitemap
Why it happens: Redirects, 404s, noindex pages and non-canonical duplicates. They waste crawl on pages that will never be indexed and bury real problems in the Search Console reports.
How to fix it: List only canonical URLs that return a 200 status and are meant to be indexed. Let your CMS generate the file so deleted pages drop out automatically.
Sitemap best practices: what to include and what to leave out
A valid sitemap is the starting point. A useful one follows a few rules:
- List every page you want indexed. Product pages, articles, categories, landing pages. Our data shows that the share of your site in the sitemap is the strongest predictor of how much of it Google indexes.
- Leave out everything else. Redirects, 404s, noindex pages, filter and parameter duplicates, and staging URLs don't belong in a sitemap.
- Use canonical, absolute URLs. One protocol, one host, and the same URL your canonical tags point to.
- Keep lastmod honest. Update it when the content changes, not on every deploy.
- Generate it automatically. From your CMS or at build time, so a new page is in the sitemap the moment it is published and a deleted one drops out.
- Split large sites by page type. Separate sitemaps for products, articles and categories let you filter the Search Console page indexing report by sitemap and spot the section Google is ignoring.
How to submit your sitemap to Google and Bing
- Declare it in robots.txt. Add
Sitemap: https://yoursite.com/sitemap.xmlso every crawler that reads robots.txt can find it. - Submit it in Google Search Console. Open the Sitemaps report, paste the URL under Add a new sitemap, and click Submit. The report then shows when Google last read the file and how many pages it discovered.
- Submit it in Bing Webmaster Tools. Bing also supports IndexNow, which notifies it of new URLs instantly. Google does not support IndexNow.
- Don't rely on pinging. Google deprecated its sitemap ping endpoint in 2023. An accurate lastmod and a resubmission in Search Console are how you signal changes now.
Where to find your sitemap on WordPress, Shopify, Wix and Next.js
Most platforms generate a sitemap for you. If you don't know the URL, enter your domain in the validator above and it will look in these places:
| Platform | Default sitemap | Notes |
|---|---|---|
| WordPress | /wp-sitemap.xml | Built in since WordPress 5.5. Yoast SEO and Rank Math replace it with /sitemap_index.xml. |
| Shopify | /sitemap.xml | Generated automatically, as an index of product, collection, page and blog sitemaps. |
| Webflow | /sitemap.xml | Switch on Auto-generate sitemap in the SEO tab of your site settings. |
| Wix | /sitemap.xml | Generated automatically by Wix. |
| Squarespace | /sitemap.xml | Generated and updated automatically. |
| Ghost | /sitemap.xml | Generated automatically, split into posts, pages, tags and authors. |
| Next.js | /sitemap.xml | Not generated by default. Add app/sitemap.ts (App Router) or build the file at deploy time, and include every dynamic route. |
How BlogSEO keeps your new pages discoverable
Publishing is only half the job if Google never finds the page. BlogSEO writes and publishes SEO articles to your CMS on a schedule, and once you connect Google Search Console, its auto-indexing resubmits your sitemap to Google every time a new article goes live. The dashboard then shows how many of your pages Google has indexed, using the same URL Inspection data behind this study. For the rest of your site, run our free Technical SEO Audit to check indexability and on-page basics, and the Google Rank Checker to see where your pages rank once they are indexed.
Access SEO tools from any page with the BlogSEO Chrome extension
Get instant SEO insights on the site you're browsing, without leaving the tab. The same free tools, one click away in your toolbar.
- 100% free, no account required
- Works on any website you visit
- Installs in seconds from the Chrome Web Store
Free forever • No signup • Takes 10 seconds
FAQ
- Do I need a sitemap if my site has good internal links?
- Does a sitemap guarantee that Google will index my pages?
- How long does it take Google to index pages from my sitemap?
- How do I make sure Google knows about my sitemap?
- Why does Search Console say my sitemap could not be read?
- What does "Discovered - currently not indexed" mean?
- What is an XML sitemap?
- What is a sitemap index?
- How large can my sitemap be?
- Does Google read my sitemap automatically if it's in robots.txt?
- Is this Google's official sitemap validator?
- Does this tool support HTML sitemaps?
- Can I validate a sitemap that requires authentication?

Free SEO playbook
Rank #1 on Google & ChatGPT in the AI era
Get Rank Like Crazy, the playbook with 10,000+ copies downloaded, plus the keyword research, AI citation templates and crawlability checklist that come with it. Free.