llms.txt Generator

Give AI a file it will actually read — this tool crawls your site, picks the pages worth listing, and shows every decision it made.

What is llms.txt?

llms.txt is a proposal for a simple convention: publish a markdown file at/llms.txt that tells language models what your site is and which pages are worth their attention. It was proposed in September 2024 and is now maintained asa v2 proposal at llmstxt.org. The file follows a fixed shape — an H1 with the site name, a blockquote summary, then markdown link sections — so both people and agents can parse it at a glance.

It sits alongside two older conventions. robots.txt answers “may you fetch this?” for crawlers; sitemap.xml is a raw, machine-oriented list of URLs. llms.txt answers a different question: “of everything here, what should you actually read?” A related but separate idea, ai.txt, explores per-site AI preferences — llms.txt is the convention that saw adoption.

The format, with a template

The proposal is strict about a few things: the H1 must be your site name, the blockquote summary is required, links use markdown with absolute URLs, and optional, lower-priority pages belong under an Optional heading. Here is a workingllms.txt template you can copy:

# Acme

> Acme builds invoicing software for independent developers.

## Core

- [Pricing](https://acme.com/pricing): Three tiers, including a free plan.
- [Features](https://acme.com/features): What the product does, with screenshots.

## Docs

- [Getting started](https://acme.com/docs/start): First invoice in five minutes.
- [API reference](https://acme.com/docs/api): REST endpoints and authentication.

## Optional

- [Changelog](https://acme.com/changelog): What shipped recently.

Every file this generator produces follows that structure — including the section naming — so you can hand-edit the output without breaking the convention.

Real llms.txt examples

Files in the wild are instructive, including where they deviate from the proposal. All four of these are live files, quoted with their real sizes:

SiteSizeWhat it shows
This site (llmstxtscan.com)~1 KBThe proposal followed exactly — it doubles as the first file this tool validated.
llmstxt.org637 bytesThe proposal’s own file: minimal, with a required blockquote.
docs.anthropic.com69,795 bytesA large docs implementation — note it skips the blockquote.
cursor.com22,362 bytesA bare link list: useful to agents, loose on format.

Who uses llms.txt?

The strongest adoption signal so far comes from AI vendors documenting their own platforms:OpenAI’s developer docs,Anthropic’s developer docs andGoogle’s Gemini API docs all publish llms.txt files. On the tooling side, Chrome Lighthouse now audits for an llms.txt file under its Discoverability checks, so a missing file shows up the next time you run an audit. WordPress plugins such as Yoast SEO have added generation support as well — for anyone not on those platforms, a standalone generator like this one fills the gap.

FAQ

What is a llms.txt file?
A markdown file published at the root of your site (e.g. https://example.com/llms.txt) that gives AI models a short summary of the site plus curated links to its most useful pages, grouped under readable headings.
How is llms.txt different from robots.txt?
robots.txt controls access: it tells crawlers what they may or may not fetch. llms.txt provides guidance: it suggests what is worth reading. One restricts, the other recommends — most sites benefit from having both.
Is llms.txt an official standard?
No. It is an open proposal, not a W3C or IETF standard. That said, major AI vendors publish llms.txt files for their own developer documentation, and Chrome Lighthouse now includes an llms.txt audit in its Discoverability checks.
Does AI actually read llms.txt?
Honestly: there is no public evidence that every AI client reads it. What is verifiable is that OpenAI, Anthropic and Google maintain llms.txt files for their own docs, and tooling support keeps growing. It costs almost nothing to add, which is why the pragmatic move is to publish one.
Which pages should I include in llms.txt?
Pages that stand on their own for someone trying to understand or use your site: pricing, product overviews, key documentation, and a few strong blog posts. Exclude utility pages (login, checkout), archives (tag and category listings), and near-duplicates. The decision table above shows exactly which pages this tool includes and why.
What is the difference between llms.txt and llms-full.txt?
llms.txt is the compact index: one line per page. llms-full.txt is the full-text companion that concatenates the complete content of every included page. Start with the compact file; add the full version when an agent needs to read everything in one request.
Where should the file be placed?
At the root of your site, reachable at /llms.txt, with absolute URLs in every link. A file anywhere else defeats the convention clients expect.
Do I need to regenerate the file after my site changes?
Yes, whenever pages are added, removed or restructured — a stale index misleads both AI and human readers. Since generation takes seconds, rerunning the tool on each release is the easy habit.
Do you store the URLs or files I submit?
No. The generation happens per request and nothing is persisted — no database, no logs of your URLs, no account.
Can I publish the generated file as-is?
You can, but read the inclusion decision table first. The rules are conservative, and you know your site better than any crawler — moving a page between sections or fixing a description takes a minute.