nalyzed.

Technical SEO checklist

By Nalyzed · Published 2026-10-07 · Updated 2026-10-07 · 8 min read

Technical SEO problems are not equally important. A noindex tag or a blocked robots.txt rule removes a site from search entirely; a missing Open Graph tag costs nothing in ranking. Work in order of severity: indexability first, then canonicalization and duplication, then crawl efficiency, then content structure, then the refinements.

1. Indexability — the failures that cost everything

Check these before anything else. Each one can remove a page, a section or an entire site from search, and each is invisible in day-to-day use.

  • No unintended noindex in a robots meta tag or an X-Robots-Tag response header. Check the header too — it is the one people forget.
  • robots.txt does not disallow pages you want indexed, and does not block the CSS and JavaScript needed to render them.
  • The production site is not still carrying a staging-era blanket block.
  • Important pages return HTTP 200, not a soft 404 that answers 200 with an error page.
  • Real 404s return 404, and real redirects return 301 or 308. Streaming frameworks commonly answer 200 with not-found content, which turns every bad URL into a thin indexable page.
  • HTTPS works, the certificate is valid for the hostname actually served, and HTTP redirects to it in one hop.

2. Canonicalization and duplication

These do not remove you from the index; they split your signals across several URLs so none of them is strong.

  • Every indexable page has a self-referential canonical tag.
  • No page pairs a canonical pointing at a different URL with a noindex directive — that is a contradictory instruction, and Google treats the canonical as a request to index the other URL instead.
  • One hostname is canonical. Pick apex or www, redirect the other, and do the same for trailing slashes and uppercase paths.
  • Parameterized URLs do not create crawlable duplicates of the same content.
  • Titles and meta descriptions are unique per page. Duplicates across a template are the usual symptom.
  • hreflang sets are reciprocal, use valid language tags, and include an x-default.

3. Crawl efficiency

Relevant once the site is bigger than a few hundred pages, and critical past a few thousand.

  • An XML sitemap exists, is declared in robots.txt, and lists only canonical, indexable, HTTP 200 URLs.
  • lastmod is accurate. A timestamp that changes on every build is worse than no lastmod at all.
  • No sitemap exceeds 50,000 URLs or 50MB uncompressed; split into a sitemap index beyond that.
  • Infinite crawl spaces are blocked — calendars, faceted filters, cursor or session parameters.
  • Redirect chains are collapsed to a single hop, and there are no redirect loops.
  • Internal links point at canonical URLs directly instead of hopping through a redirect.
  • No orphan pages: every page you want indexed is linked from at least one other indexed page.

4. Content and structure

  • One H1 per page, with a heading hierarchy that does not skip levels.
  • The substance of the page is in the server-rendered HTML, not only assembled client-side.
  • Titles are descriptive and distinct, roughly 15 to 65 characters so they are not truncated.
  • Meta descriptions are written per page, roughly 50 to 160 characters. They are not a ranking factor but they drive click-through.
  • Images have meaningful alt text, and explicit width and height so they do not shift layout.
  • Structured data is valid, matches the visible content, and names a publisher. Article markup needs a headline, a date and an author.

5. The refinements

Worth doing, but only after everything above holds. None of these will rescue a site with an indexability problem.

  • Open Graph and Twitter Card tags for link previews — these affect sharing, not ranking.
  • max-snippet:-1 and max-image-preview:large on pages whose value is being quoted.
  • Security headers: HSTS, Content-Security-Policy, X-Content-Type-Options, Referrer-Policy.
  • Core Web Vitals, measured in the field rather than in a single lab run.
  • An IndexNow submission on publish, which gets new pages recrawled by Bing and Yandex without waiting for sitemap discovery.

How to work through it

Run the list top to bottom and stop at the first category with failures. Fixing a canonical problem while a noindex tag is still live accomplishes nothing, and tuning Core Web Vitals on a page nobody can crawl accomplishes less.

Re-check after each deploy. Most of these regress silently — a header changes, a template ships a stray noindex, a sitemap starts listing redirects — and nothing in a normal test suite notices.

Frequently asked questions

What is the single most common technical SEO mistake?
An unintended noindex surviving into production, usually in an X-Robots-Tag header rather than the HTML, where it is easy to miss. The second most common is a soft 404: a missing page answering HTTP 200 with error content, which search engines treat as thin content rather than a removal.
Do I need a sitemap if my site is small?
Not strictly — a small, well-linked site will be crawled without one. It is still worth having: it is cheap, it communicates lastmod, and it gives you a coverage report to compare against in Search Console.
Does duplicate content get a site penalized?
There is no duplicate content penalty. What happens is signal dilution: search engines pick one URL to index and your links and authority are spread across the variants. Canonical tags and consistent internal linking are how you consolidate that.
How often should I run a technical audit?
Re-run the indexability checks on every significant deploy, since those regress silently and cost the most. A fuller pass each quarter, or after any template, platform or hosting change, is usually enough.

Check your own site

Published by Nalyzed. Scoring rules and limits are documented on the methodology page.