nalyzed.

GEO readiness: making a site AI search engines can cite

By Nalyzed · Published 2026-10-07 · Updated 2026-10-07 · 7 min read

Generative engine optimization (GEO) is the practice of making a page easy for an AI answer engine to fetch, understand and attribute. Unlike ranking, none of it can be verified from the outside — so the only honest approach is to optimize the four things that are measurable: whether answer-engine crawlers can reach the page, whether the answer is in visible HTML text, whether the page carries attribution cues such as an author, a date and outbound sources, and whether the page is discoverable at all.

GEO is not a ranking you can check

There is no GEO equivalent of a rank tracker. Answer engines do not publish a position for your page, citations vary between identical prompts, and the same question asked twice can produce different sources. Any tool claiming to measure your AI ranking is inferring it from a handful of sampled prompts.

What can be measured is readiness: the technical and editorial conditions that have to hold before a page can be cited at all. That is what Nalyzed's GEO readiness score reports, and why it is deliberately separate from the composite score — it is a diagnostic, not a prediction.

1. Crawler access decides everything else

An answer engine that cannot fetch your page cannot cite it. This is the single most common and most invisible failure: a robots.txt line added years ago, or a blanket block copied from a template, quietly removes a site from every AI surface.

Separate the two decisions. Training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended) collect data to train models; blocking them is a legitimate choice with no effect on whether you are cited. Search and user crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User) are what fetch pages to answer a live question. Blocking those removes you from the answer, not from the training set.

  • Check robots.txt per crawler rather than assuming a wildcard rule covers them.
  • Watch for blank lines and comments inside a group — many parsers, and plenty of hand-written files, get this wrong.
  • Do not send nosnippet or max-snippet:0 if you want to be quoted: they explicitly forbid the excerpt an answer needs.
  • Prefer max-snippet:-1 and max-image-preview:large on pages whose value is being quoted.

2. The answer has to be in the HTML

Answer engines generally do not execute JavaScript the way a browser does. If the substance of your page arrives via client-side rendering, a crawler may see an empty shell — and a shell cannot be summarized.

Structure matters as much as presence. A page that states its conclusion in the first paragraph, then supports it under descriptive headings, is far easier to extract a correct answer from than the same information buried after a hero section and three carousels.

  • Lead with the answer. Put the conclusion in the first paragraph, before the context and the caveats.
  • Use headings that match real questions, not marketing labels.
  • Keep the key facts in text. A number that only exists inside a chart image is invisible.
  • Avoid hiding substance behind tabs or accordions where you can; collapsed content is readable but weighted less.

3. Attribution cues make a page safe to quote

An engine deciding whether to cite a page is implicitly judging whether the citation will embarrass it. Pages with no author, no date and no sources look unattributable, and unattributable pages make poor citations regardless of how good the content is.

This is the cheapest category to fix and the most commonly skipped: a named author, a visible last-updated date, outbound links to primary sources, and schema.org markup naming the publisher.

  • Name a real author or a named organization, in visible text as well as in markup.
  • Show a publication and last-updated date, and keep it honest — a date that changes on every deploy trains crawlers to ignore it.
  • Link to the primary sources behind your claims.
  • Declare a publisher with schema.org Organization markup, referenced from each page.

4. Discovery still works the old way

Answer engines find pages largely the way search engines do: sitemaps, internal links and existing index coverage. A page nobody links to and no sitemap lists is unlikely to be reached by anything.

An llms.txt file is a cheap, optional extra on top of this — a curated map of your most useful pages. Support is still emerging and it is not a ranking factor, but it costs almost nothing to publish.

  1. Make sure the page is in an XML sitemap that robots.txt actually declares.
  2. Link to it from at least one page that is already indexed.
  3. Keep the canonical URL self-referential, and never pair a canonical pointing elsewhere with a noindex directive.
  4. Optionally publish /llms.txt listing your highest-value pages.

What to do first

In order of impact: confirm search-and-user crawlers are not blocked, confirm the substance of the page is in the server-rendered HTML, add an author and a last-updated date, then worry about llms.txt and structured data.

The first two are binary failures — they either work or the page is invisible. The rest are improvements on a page that already qualifies.

Frequently asked questions

Is GEO different from SEO?
It overlaps heavily. Crawlability, server-rendered content, internal linking and structured data serve both. GEO adds emphasis on answer-first writing, explicit attribution cues and per-crawler access control for answer engines, and removes the parts of SEO that are about competing for a ranked position.
Does blocking GPTBot hurt my visibility in ChatGPT?
No. GPTBot collects training data. OAI-SearchBot and ChatGPT-User are what fetch pages to answer a live question. Blocking GPTBot while allowing those keeps you citable while opting out of training.
Does llms.txt improve my AI visibility?
There is no evidence it acts as a ranking or citation factor, and support varies by tool. It is worth publishing because it is nearly free and helps anything that does look for it, but it will not compensate for a page an engine cannot fetch or read.
Can I measure whether AI engines cite my site?
Only by sampling: asking engines questions you expect to rank for and recording whether you appear. That is useful directional feedback, not a metric. Readiness signals are what you can measure reliably and act on.

Check your own site

Published by Nalyzed. Scoring rules and limits are documented on the methodology page.