Home › Guides › Heading Structure in the Wild: 53 Homepages

Heading Structure in the Wild: 53 Homepages

We parsed 1,696 headings from 53 well-known homepages: 16 have no single visible H1, and 22 open with something other than an H1.

How this was measured

On 2026-10-03 we requested the homepage of each of the same 78 hand-picked domains our other surveys use — news, commerce, SaaS, developer tooling, social, streaming, reference, SEO, education, finance, cloud, local and government. We followed redirects, because the point is to read the page a site calls its homepage, and eight of these domains do not serve it at /. Every request identified itself with a survey user agent and was attempted directly first, then through a local proxy; 52 domains answered on the direct path and 24 needed the proxy.

53 of 78 produced a readable homepage (67.9%). The other 25: 15 answered with a bot wall (403, 401, 406 or 202), 3 served a challenge page with a 200 — archive.org returned a 1,872-byte "Javascript is required" shell, khanacademy.org a "Client Challenge" shell and paypal.com a "Please wait while we perform security check" page — 5 returned a 200 with no headings in the HTML at all, and 2 never answered (bbc.com, washingtonpost.com). The challenge pages are excluded rather than counted: a captcha with no H1 is not a finding about heading structure.

The 53 homepages yielded 1,696 headings. We blanked script, style, template and noscript before parsing, replacing them with equal-length spaces — without that step, a framework's client-side template contributes headings that were never in the DOM. We matched nav, header, footer, aside, main and form with a proper open/close tag stack, so a heading is attributed to its innermost container rather than to the nearest preceding tag. Hidden headings were judged conservatively: only an explicit aria-hidden="true", a hidden attribute, a screen-reader class (sr-only, visually-hidden, screen-reader-text) or an inline display:none counts. Anything that requires reading a stylesheet is invisible to us, so our hidden count is a floor, not a total.

Finding 1: one visible H1 is the majority, not the norm

37 of 53 (69.8%) have exactly one visible H1. That sounds reassuring until you look at the other 16, which split three ways. 5 sites ship two or more: x.com (2: "Happening now." and "See what's happening"), asana.com (2), cloudflare.com (2), linkedin.com (5) and python.org (5). 11 have no visible H1 at all.

None of the multi-H1 sites repeated the same text, so these are not copy-paste accidents. They are pages built as a stack of independent sections, each written by a different team, each with its own top-level heading. python.org is the clearest case: its five H1s are the five slides of its hero carousel — "Intuitive Interpretation", "Compound Data Types", "All the Flow You'd Expect", "Functions Defined", "Quick & Easy to Learn" — and all five sit inside a single <header class="main-header"> element. It is a defensible piece of HTML in isolation; it is not a page with a topic.

Finding 2: the 11 missing H1s are two completely different problems

Eight have no <h1> element anywhere: theguardian.com, cnn.com, forbes.com, bloomberg.com, trello.com, pinterest.com, spotify.com and cdc.gov. That is all four news sites in our sample — every one of them. A news homepage is an index of links, and none of the four chose to name it.

Three do have an H1, hidden on purpose: mit.edu ("Massachusetts Institute of Technology"), harvard.edu ("Harvard University") and stanford.edu ("Stanford University"), each with a screen-reader class. This is the correct accessible pattern for a page whose visual design has no room for a visible title: the document still has a name, assistive technology can announce it, and nothing looks wrong. It is the opposite of a defect, and it is why "no visible H1" should never be reported without checking whether an H1 exists at all.

Two of the eight are worth naming separately. cnn.com's first heading is an H3 inside a <form> reading "CNN values your feedback" — the top of CNN's heading outline belongs to a survey widget. pinterest.com has exactly one heading on its entire homepage: an H2 reading "Sign up to get your ideas", which is the logged-out gate, not the page.

Finding 3: 22 homepages do not open with an H1

The first heading in document order is an H1 on 31 of 53 (58.5%). On the other 22 it is lower: 10 start at H2 (github.com, atlassian.com, figma.com, pinterest.com, spotify.com, mozilla.org, w3.org, stanford.edu, netlify.com, cdc.gov), 9 start at H3 (theguardian.com, cnn.com, forbes.com, bloomberg.com, trello.com, semrush.com, searchenginejournal.com, digitalocean.com, heroku.com) and 3 start at H4 (kubernetes.io, moz.com, squareup.com).

Most of these are navigation and section labels arriving before any content, which is the ordinary consequence of a masthead built from headings. Three are structural choices instead: bloomberg.com has 66 headings and not a single H2 — its navigation columns are H3s and nothing above them; trello.com uses H3 as its working top level throughout; and nodejs.org, on 531 KB of HTML, has exactly one heading on the page, its H1 "Run JavaScript Everywhere".

Finding 4: depth is decoration, and H3 outnumbers H2

Across all 53 homepages the level distribution is H1 ×57, H2 ×640, H3 ×793, H4 ×119, H5 ×47, H6 ×40. There are more H3s than H2s. That is what happens when a card grid is built one level deeper than its section heading, repeated across dozens of cards — the level was chosen for the font size it produces, not for the outline it implies.

11 sites (20.8%) skip a level somewhere in their visible outline: moz.com six times, asana.com and cloudflare.com five each, wise.com three, docker.com two, and kubernetes.io, netflix.com, vimeo.com, screamingfrog.co.uk, mit.edu and squareup.com once each. 16 sites go as deep as H4 or beyond, and three reach H6 (asana.com, squareup.com, cloudflare.com). A homepage with six levels of heading is a homepage whose headings stopped meaning anything.

Finding 5: a fifth of all headings belong to the furniture

364 of 1,696 headings (21.5%) sit inside nav, header, footer or aside: 160 in nav, 99 in header, 92 in footer and 13 in aside. These are headings doing the job of a label — "Site navigation", "Popular artists", "Information for" — and they land in the same outline as the headings that describe the page.

The practical consequence is that an outline read top to bottom is not a summary of the page. stanford.edu opens with five H2s, all inside nav: "Our Impact", "About Stanford", "Information for", "Academics Overview", "Information for". Read as content, that is nonsense; read as navigation, it is fine. Anything you feed to a crawler or an assistant that flattens headings into an outline will show these first.

Finding 6: 19 hidden headings, and one that should not be hidden

10 sites hide 19 headings. Most are correct: netlify.com hides an H2 reading "Site navigation", and MIT, Harvard and Stanford hide their H1s, as described above. vimeo.com hides 6, github.com 3 and coursera.org 3.

One is a bug. stripe.com renders its hero H1 — the one visible sentence describing the page — with aria-hidden="true". The heading is on screen and has been removed from the accessibility tree, so a screen reader reaches the page's main title and finds nothing. That is the reverse of the universities' pattern and the one thing on this list worth copying nowhere.

Finding 7: 17 empty headings, mostly placeholders

Five sites ship 17 headings with no text content: digitalocean.com (5), target.com (4), squareup.com (4), who.int (3) and stripe.com (1). The common case is an sr-only H3 whose text is injected later — target.com has four H3s reading "Loading..." — which is a heading element used as a live region. It is invisible in rendering and meaningless in an outline, but it is still a heading as far as any parser is concerned.

Finding 8: when the H1 exists but says nothing

Two sites write an H1 that is technically present and semantically empty: target.com uses "Homepage" and airbnb.com uses "Airbnb homepage". Both are label-position text — the word that fills the space where a title goes — rather than a description of what the page is.

At the other extreme, screamingfrog.co.uk writes a 104-character H1: "We are a UK based SEO agency and the creators of the famous website crawler and log file analyser". An H1 is not a meta description; if it takes three lines to read, it is doing the wrong job.

H1 and <title> match exactly once in 53 — walmart.com, where both read "Walmart | Save Money. Live better." A 1.9% match rate is not a problem to fix; the two elements have different jobs, and only one of them has to survive being read aloud by a screen reader.

Scale: what a homepage looks like

The median homepage has 27 headings and a median of 9 H2s. The widest is theguardian.com at 196 headings (33 of them H2), followed by searchenginejournal.com at 106 — of which 92 are H2s, a blog index where every post title is a top-level section. The narrowest are pinterest.com and nodejs.org at one heading each, then x.com at 2, airbnb.com at 4 and linkedin.com at 5.

By category, the developer sites are the most conventional (6 of 7 have exactly one visible H1) and so are the SEO sites (6 of 6). Education is the outlier in the other direction: 1 of 4, because MIT, Harvard and Stanford all hide theirs.

What to do with this

Check whether you have an H1 before you check whether it is visible. Eight of the 53 sites here have no H1 element at all; three have one that is hidden on purpose. Those need opposite responses, and a tool that only counts visible headings cannot tell them apart. Our heading structure analyzer reports both the count and what each heading says, which is the only way to distinguish them.

Write the H1 as the name of the page, not the name of the site. "Homepage" and "Airbnb homepage" are the failure mode. If you cannot write a sentence, the honest answer is that the page has no single topic — which is true of the four news homepages here, and is why none of them wrote one.

Do not use heading levels for font size. The H3-over-H2 imbalance and 11 sites with level skips both come from styling. Pick the level that describes the relationship, then style it with a class. If you need H6, you need a shorter page.

Keep navigation out of the heading outline. 364 headings here belong to nav, header, footer or aside. A footer link column does not need to be an H3; a <ul> with a visually hidden label does the same job without entering the outline.

Never put aria-hidden on visible text. One site in 53 does it to its own H1. If a heading is decorative, it should not be a heading. If it is your page title, it must stay in the accessibility tree — hide it with a screen-reader class, the way MIT, Harvard and Stanford do, not with aria-hidden.

Then check what the page looks like to something that cannot run JavaScript. Four of the five homepages we could not read — Reddit, YouTube, Tumblr and Twitch — return server HTML with no headings at all. Google renders JavaScript, so those pages are not invisible to it, but they are invisible to every simpler consumer, and you cannot tell from the rendered page alone.

Limitations

53 homepages out of 78 attempted, from a hand-picked list of well-known sites, not a random sample of the web. Categories are our own manual grouping; several contain three or four domains, and news is represented entirely by four sites that happen to share one design pattern.

The largest limitation is JavaScript. We read server HTML only. Four of the five pages that returned 200 with no headings — reddit.com, youtube.com, tumblr.com and twitch.tv — render their content client-side, so we recorded them as unreadable rather than as pages without headings. wikimedia.org is the fifth and is genuinely heading-free in its server HTML. Anything we say about heading counts describes what a non-rendering client sees, which is not what Google sees.

15 domains answered with a bot wall and 3 with a challenge page. Those are excluded rather than counted, and a browser-like user agent would have reached more of them. Our hidden-heading detection reads only inline attributes and well-known class names, so it undercounts: a heading hidden by an external stylesheet is counted as visible. Empty headings are counted as headings, which inflates totals for the five sites that use them as loading placeholders.

Requests came from a single survey user agent on one day over two egress addresses, one of them in China. Several of these sites answered with a regional edition — linkedin.com redirected us to linkedin.cn and served five Chinese-language H1s, and our redirect chain survey found the same geo behaviour on Stripe, PayPal, Netflix and Square. What we parsed is the edition we were served.

Container attribution relies on balanced tags. We verified the ones we quote — python.org's <header> spans bytes 6,876 to 28,791 and all five of its H1s fall between 21,011 and 28,076 — but a site with an unclosed <nav> would be mis-attributed.

We measured structure, not quality. A page with one H1 and nine H2s may still say nothing, and the outline is a proxy for how a page is organised, not a judgement of whether it is any good.

Reproduce it

The script and the shared domain list ship with this site. node scripts/survey-headings.mjs fetches all 78 homepages and stores every response; --report re-derives every number from stored data without touching the network; --evidence prints the raw tag and inner HTML behind each verdict, which is how every heading quoted above was verified; --detail prints the per-level distribution and the extremes. Our robots.txt, canonical, hreflang, AI crawler, llms.txt, sitemap and redirect chain surveys run the same domain list from other angles.

About the author

Hongtao Ren (任宏涛) — Developer based in Xi'an, China. Builds browser-based tools and JetBrains IDE plugins. He built and maintains SerpPrism.

Corrections are the most useful thing you can send. If a tool or guide here gives you a wrong answer, that is a bug, not a judgement call — use the contact page.

Questions

How many H1s should a page have?

One, describing what the page is. 37 of the 53 homepages we parsed have exactly one visible H1; 5 have two or more and 11 have none visible — though 3 of those 11 do have an H1, hidden with a screen-reader class on purpose.

Is a hidden H1 bad for SEO?

No, and it is often the correct accessible pattern. MIT, Harvard and Stanford hide theirs with an sr-only class so assistive technology can announce a page name the visual design has no room for. What is wrong is aria-hidden on a visible heading, which one site here does to its own H1.

Does Google penalise multiple H1s?

No. Google has said it does not, and five of the homepages here ship two or more. The cost is ambiguity rather than penalty: with five H1s there is no single answer to what the page is about, which matters more for AI summaries and assistive technology than for rankings.

Do heading levels affect rankings?

Not as a ranking factor. They affect how a page is understood: 11 of the 53 sites here skip a level and 16 go as deep as H4, and in most cases the level was picked for the font size it produces. Use the level that describes the relationship and style it with a class.

Should footer links be headings?

Usually not. 364 of the 1,696 headings we counted — 21.5% — sit inside nav, header, footer or aside. A footer column does not need an H3; a list plus a visually hidden label keeps it out of the outline.

Why do news homepages have no H1?

Because they are indexes of links rather than pages with a topic. All four news sites in our sample — the Guardian, CNN, Forbes and Bloomberg — ship no H1 element at all. Whether that is a mistake depends on whether you think the page has a name.

How do I check my own page?

Paste the HTML into the heading structure analyzer and read the outline rather than the warning count. Check three things in order: does an H1 element exist, is it visible, and does it name the page rather than the site.