You can tell if someone has been doing SEO for a while by whether or not they know what an HTML sitemap is. In the past, creating an HTML sitemap was an ordinary step in building a website. It didn’t take much thought or effort, so everyone created them.
When XML sitemaps came along, however, most people abandoned HTML sitemaps. However, a recent discussion between Google’s John Mueller and Martin Splitt indicated that maybe people gave up on HTML sitemaps too soon.
In a recent episode of Google’s Search Off The Record podcast, Splitt asked a question that continues to puzzle many site owners today: do we have to use XML sitemaps, or can we also use HTML sitemaps? Mueller answered first with some of the same confusion that has plagued many of my clients for years.
“Yeah, I find that many people are confused about how HTML sitemaps relate to their overall website development. Most individuals understand what XML and HTML are; however, when they look at an XML document, they often confuse it with HTML. They will typically see the strange-looking tags and assume “this is HTML,” or “it looks a lot like HTML.”
This resembles an odd programming language. The two types are referred to as “the same thing.” However, an HTML sitemap is essentially like a “map” of your website for users. This is not meant to take the place of an XML sitemap file.
If a person is scanning your site and locates an HTML sitemap file, then they may use the same method of scanning to scan the links contained in the sitemap file as they would for any other link within your site. Therefore, locating the HTML sitemap file would be beneficial to the scanner.
But you can’t use the HTML sitemap in place of the XML sitemap. You can’t submit an HTML sitemap to Google and expect them to process it in the exact same manner as they do an XML sitemap. An HTML sitemap is merely a list of links.
I have had the same discussion with many of my non-technical clients. Although both types of sitemaps have the same name, they are very different.
An XML sitemap is a machine-readable document that has a specific format. It is created so that you can submit it to search engines. On the other hand, an HTML sitemap is a regular web page that contains many links.
HTML sitemaps were developed to help people find information on your site. As long as Google crawls the HTML sitemap in the same manner as it does other pages on your site, Google will be able to index the content of the HTML sitemap. However, Google cannot process the HTML sitemap as a submission of a sitemap.
HTML sitemaps were first used to provide users with a visual representation of the structure of their website. In addition to providing a clear picture of how a user would navigate through a site, HTML sitemaps also provided search engine bots with an easy way to understand the relationships between pages on a site.
As time went by, HTML sitemaps became less common as most sites moved towards using XML sitemaps as a primary method for communicating site architecture to search engine spiders. XML Sitemaps differ from HTML sitemaps in that they are created in XML format, which makes them much easier for search engine spiders to parse than HTML-based sitemaps. Also, unlike HTML sitemaps, XML sitemaps can be submitted directly to major search engines (like Google) so that the search engine can begin indexing the pages listed in the XML sitemap immediately.
The reason that this is so important is that it provides some perspective as to what the indexing process used to look like prior to the introduction of XML sitemaps. Prior to XML sitemaps, the only way for search engines to find a new page on your site was through links (either from your own pages or from other sites).
In addition to linking, many people would include a link in the footer of their home page or on each page of their site pointing to an HTML sitemap. This HTML sitemap was essentially a long list of links to the internal pages that you wanted the search engines to crawl and rank.
While this was a very basic process, it was effective during a period when search engine spiders were not as advanced in discovering content.
The search engines were always looking for new methods to find fresh content. RSS feeds would have been a great way to do so. In 2004, a university study showed that crawlers could utilize RSS feeds to identify new and recently updated pages, reducing the crawl bandwidth used by as much as 40 percent.
However, since RSS was never fully accepted by all users, it did not provide the reliability needed to serve as a discovery tool. This lack of adoption is most likely what drove Google, Yahoo, and Microsoft to agree on a single XML sitemap standard that remains the current industry standard.
So Why Would Anyone Still Build One
The thing I found most surprising was that Mueller did not say, “This is an old-fashioned way of doing things.” He said, “An HTML sitemap is a way to help people navigate through your website.” He used an e-commerce site to illustrate this point, but it could apply to any website. A well-designed HTML sitemap will allow users to get to the main categories of your website quickly, and then to find the exact item or subject matter that they are looking for.
The question posed by Splitt connects this to the concept of crawling.
“The other way of doing things may also assist with crawling since there are many links to be found within this location; however, I believe it isn’t quite as organized or clearly defined as an XML sitemap.”
Mueller stated this, and additionally provided information regarding what content is expected to be included on each page
“Yes, definitely. I believe many HTML sitemaps do not contain all the content of the site.”
“In other words, you will not list each of your products on your e-commerce site within a single HTML sitemap. Instead, you may create a list of product categories that users can click upon to find the specific product they are looking for. However, once users click on a category, they will need to select the individual product. This is similar to what we discussed earlier, however, it is slightly different.”
A modern-day HTML sitemap differs from the previous “dump all of the URLs” method most of us used several years ago.
Today’s HTML sitemap provides a user-friendly, curated view of how your website is structured. It does not provide a crawl budget trick.
I’ve been working in the industry for 20 years now, and I wanted to give you my opinion on the state of things.
This idea may be new to some readers. To other readers, however, they will group this into the category of “directory submissions” or “reciprocal link swaps,” which was popular back in 2005.
I can see why many people would be skeptical. After all, there are many “skeletons” in the SEO closet. However, John Mueller is not suggesting that you use a strategy from 2005. He is stating that the page can assist users, and anything that assists users with navigation on your website will also assist Google with understanding your website.
Things To Keep In Mind If You Choose To Create HTML Sitemap
Design the site for real people as much as possible. Make sure the organization reflects what a real person would want to find on the site, with logical groupings and descriptive labels.
Be cautious. When linking from a large website, only link to the main category or key hub page, do not link to each URL.
Keep XML as well. Your HTML version will supplement your XML sitemap, not replace it, and Mueller has been very clear about this.
Link to the URL in an area that is easily seen; the footer has been the standard location for links to the URL, and it continues to be effective.
The more pages a visitor can navigate through on your website, the longer they will remain on your site; the greater the number of other pages a visitor explores and the larger the amount of content a visitor shares. All three of these actions provide positive feedback to Google, which is then able to recognize them.
The HTML sitemap, while simple and inexpensive to implement, is a powerful tool for improving both the ability to locate specific content and the ability to find new content.