Free tool

Sitemap.xml Generator

Enter a domain. Our crawler walks the site from the homepage, one request per second, respecting robots.txt, and lists what it found as a sitemap.xml.

Up to 60 pages, one request per second, robots.txt respected, same domain only. Usually one to two minutes; you get a link to the result. Three crawls per hour per address.

About the Sitemap.xml Generator

A sitemap.xml is a list of the pages on a site that you want search engines to know about, with an optional date each was last changed. Search engines discover pages through links too, so a sitemap is not required, but it makes the set of pages explicit and is the file the audit's rule S4 looks for, at the URL named in robots.txt or at /sitemap.xml.

This generator uses the same crawler as the audit. Starting from the homepage it reads robots.txt, then follows internal links breadth-first: at most 60 pages, one request per second, no redirects to other domains, no paths robots.txt disallows for EchoRankBot. The crawl runs as a background job and usually takes one to two minutes; the result page lists every page that answered 200 with HTML, the Last-Modified date when the server sent one, and a download button for the sitemap.xml. Only the URLs are stored, for seven days, so you can come back to the result; page bodies are never kept.

Sixty pages is a sample, not a limit on your site. If the crawl reports that more pages were discovered than fetched, use the result as a starting point and add the remaining URLs from your CMS, or let the CMS generate the sitemap itself, which most platforms can do. Upload the file to the root of the site, add a Sitemap line to robots.txt with the Robots.txt Generator, and run the free score: S4 will read the live file and count its URLs.

Questions

Why did the crawl find fewer pages than my site has?

The crawler follows HTML links from the homepage only, stops at 60 pages, and skips paths robots.txt disallows. Pages reachable only through menus built by JavaScript, forms or search are not found.

What is lastmod and why is it missing on some pages?

lastmod is the date a page last changed. The tool fills it from the Last-Modified header when the server sends one; many CMS platforms do not, and the field is left out rather than guessed.

How often can I run it?

Three crawls per hour per network address. Each crawl is at most 60 requests to the site at one request per second, so it is light on the server.

What do you keep?

The list of URLs and dates, for seven days, under a link only you have. No page content is stored. The job row is deleted after that.