Where llms.txt came from
Jeremy Howard proposed llms.txt on September 3, 2024, at llmstxt.org, and has since revised it into a second version. The idea: agents now read websites constantly, a coding agent checking documentation or a chat assistant researching a product, and web pages are built for people. Navigation, ads and scripts get in the way, and most sites are too big to fit in a model’s context.
llms.txt borrows the idea of robots.txt and sitemap.xml: one file at a predictable path. It holds a short summary of the site and links to the pages that matter, written in markdown so a model can read it without an HTML pipeline.
The format
An llms.txt file is markdown with a fixed order:
- An H1 with the name of the site or project. This is the only required section.
- A blockquote summary with the key information a reader needs for the rest of the file.
- Optional detail in paragraphs or lists, without headings.
- H2 sections of links. Each item is a markdown link, optionally followed by a colon and a note on what the page holds.
- An “Optional” section, by convention, for links an agent can skip when it needs a shorter context.
The file usually sits at the root, at /llms.txt, and version 2 of the proposal also allows one at any path, such as /docs/llms.txt, covering the pages beneath it. It also proposes a markdown version of each page at the same address with .md added.
What llms.txt does and does not do
The proposal is aimed at agents reading a site while they help someone, and it says the files are used most heavily for software documentation, where coding agents follow them to reference pages. Its second version reports that thousands of sites publish one and that OpenAI, Anthropic and Gemini publish them for their own developer docs.
Search is a different matter. Google says you don’t need to create new machine readable files, AI text files, or markup to appear in AI Overviews or AI Mode. The file also has no role in access control: whether a crawler may read your site is set in robots.txt, including tokens such as Google-Extended for some of Google’s AI products.
So treat llms.txt as a small, cheap addition that helps agents that read it, and measure the AI answers you care about directly.
How it relates to robots.txt and sitemap.xml
The three files answer three different questions:
- robots.txt is access control. It tells automated tools which parts of the site they may crawl.
- sitemap.xml is an inventory. It lists every page for search engines, with no view on which ones matter most.
- llms.txt is a reading list. It says which pages to read first, with a line on each, in a file small enough to fit in context.
A site can serve all three, and each does a job the others don’t.
How to create an llms.txt file
By hand: write an H1 with your site’s name, a one or two sentence summary as a blockquote, then your most useful pages under H2 sections, each with a short note. Save it as plain text at /llms.txt and update it when your key pages change. The proposal suggests testing the file by asking an agent questions about your site with only the llms.txt to go on.
With the free llms.txt generator: it reads your sitemap, or the links on your homepage, takes the titles and descriptions of your main pages and drafts a file in the right format for you to edit and upload. No signup.