llms.txt explained
By Nalyzed · Published 2026-10-07 · Updated 2026-10-07 · 5 min read
llms.txt is a proposed Markdown file at the root of a site that gives language models a short, curated map of its most useful pages. It is not a permission file like robots.txt and not an exhaustive index like sitemap.xml — it is an editorial summary. Support is still emerging and it is not a ranking factor, but it is cheap to publish and helps any tool that looks for it.
What the file is
llms.txt is served as plain text or Markdown at the root of a domain, for example https://example.com/llms.txt. It is Markdown with a defined shape: a single H1 with the site or project name, an optional blockquote summary, optional free prose, then H2 sections containing lists of Markdown links, each optionally annotated with a short note.
The intent is curation. A sitemap lists everything you have; llms.txt lists what you would point someone at if they had thirty seconds.
How it differs from robots.txt and sitemap.xml
robots.txt grants or withholds access. sitemap.xml enumerates URLs for discovery. llms.txt does neither — it expresses priority and meaning, in prose a model can read.
That also means it carries no enforcement. Publishing llms.txt does not restrict anything, and omitting a page from it does not hide that page.
- robots.txt — machine-readable access rules, enforced by well-behaved crawlers.
- sitemap.xml — complete URL inventory with change metadata, for discovery.
- llms.txt — a short curated index with human-readable notes, for comprehension.
Writing a useful one
The common failure is treating it as another sitemap: a few hundred links with no notes, which is less useful than nothing because it buries the pages that matter.
- Start with # followed by the site or product name.
- Add a one-sentence blockquote summary saying what the site is for.
- Add a short paragraph of context if the summary needs qualification.
- Group links under H2 sections that reflect how people actually use the site — Docs, Reference, Pricing, Guides.
- Keep each section to the links you would genuinely recommend, and annotate each with a note after a colon.
- Serve it as text/plain or text/markdown, not as an HTML page.
The mistakes that make it useless
- Serving an HTML error page at /llms.txt — a soft 404 that checkers count as a malformed file rather than a missing one.
- Serving it with a text/html content type, so tools treat it as a page rather than a document.
- Listing every URL on the site, which defeats the point of curation.
- Omitting the H1, which is the only genuinely required element.
- Letting it go stale: links that 404 are worse than absent ones.
Should you bother?
Publish one if you have documentation, guides or reference material where pointing a model at the right page materially improves its answer. The cost is one file and a few minutes.
Do not expect it to move anything on its own. If answer engines cannot fetch your pages, or the substance of them is rendered client-side, llms.txt changes nothing — fix those first.
Frequently asked questions
- Is llms.txt an official standard?
- No. It is a proposal published at llmstxt.org that several tools and sites have adopted. It is not ratified by any standards body and no major engine has committed to honouring it.
- Where exactly should the file go?
- At the root of the domain: https://example.com/llms.txt, served as plain text or Markdown. Some sites also publish an expanded /llms-full.txt containing the full reference content rather than only links.
- Does llms.txt block AI crawlers?
- No. It grants nothing and forbids nothing. Access control belongs in robots.txt, and snippet control in robots meta tags or the X-Robots-Tag header.
- What is llms-full.txt?
- An informal companion convention: the same index plus the actual reference content inline, so a model can read the material without following links. It is useful for documentation sites and unnecessary for most others.
Check your own site
More guides
Published by Nalyzed. Scoring rules and limits are documented on the methodology page.