Skip to content

Glossary

llms.txt

llms.txt is a proposed markdown file, served at /llms.txt, that gives AI agents and language models a short, curated map of a website's most important content.

Also known as LLMs.txt, llms-full.txt, the llms.txt standard

Where llms.txt came from

Jeremy Howard proposed llms.txt on September 3, 2024, at llmstxt.org, and has since revised it into a second version. The idea: agents now read websites constantly, a coding agent checking documentation or a chat assistant researching a product, and web pages are built for people. Navigation, ads and scripts get in the way, and most sites are too big to fit in a model’s context.

llms.txt borrows the idea of robots.txt and sitemap.xml: one file at a predictable path. It holds a short summary of the site and links to the pages that matter, written in markdown so a model can read it without an HTML pipeline.

The format

An llms.txt file is markdown with a fixed order:

  • An H1 with the name of the site or project. This is the only required section.
  • A blockquote summary with the key information a reader needs for the rest of the file.
  • Optional detail in paragraphs or lists, without headings.
  • H2 sections of links. Each item is a markdown link, optionally followed by a colon and a note on what the page holds.
  • An “Optional” section, by convention, for links an agent can skip when it needs a shorter context.

The file usually sits at the root, at /llms.txt, and version 2 of the proposal also allows one at any path, such as /docs/llms.txt, covering the pages beneath it. It also proposes a markdown version of each page at the same address with .md added.

What llms.txt does and does not do

The proposal is aimed at agents reading a site while they help someone, and it says the files are used most heavily for software documentation, where coding agents follow them to reference pages. Its second version reports that thousands of sites publish one and that OpenAI, Anthropic and Gemini publish them for their own developer docs.

Search is a different matter. Google says you don’t need to create new machine readable files, AI text files, or markup to appear in AI Overviews or AI Mode. The file also has no role in access control: whether a crawler may read your site is set in robots.txt, including tokens such as Google-Extended for some of Google’s AI products.

So treat llms.txt as a small, cheap addition that helps agents that read it, and measure the AI answers you care about directly.

How it relates to robots.txt and sitemap.xml

The three files answer three different questions:

  • robots.txt is access control. It tells automated tools which parts of the site they may crawl.
  • sitemap.xml is an inventory. It lists every page for search engines, with no view on which ones matter most.
  • llms.txt is a reading list. It says which pages to read first, with a line on each, in a file small enough to fit in context.

A site can serve all three, and each does a job the others don’t.

How to create an llms.txt file

By hand: write an H1 with your site’s name, a one or two sentence summary as a blockquote, then your most useful pages under H2 sections, each with a short note. Save it as plain text at /llms.txt and update it when your key pages change. The proposal suggests testing the file by asking an agent questions about your site with only the llms.txt to go on.

With the free llms.txt generator: it reads your sitemap, or the links on your homepage, takes the titles and descriptions of your main pages and drafts a file in the right format for you to edit and upload. No signup.

Sources

Read on .

Frequently asked

Do ChatGPT, Google, or Claude actually read llms.txt?
Google says you don't need AI text files to appear in AI Overviews or AI Mode. The proposal is aimed at agents that read a site while helping someone, such as coding assistants reading documentation, and its second version reports that OpenAI, Anthropic and Gemini publish llms.txt files for their own developer docs. Whether a given assistant reads yours is up to the company that runs it.
Does llms.txt stop AI companies from training on my content?
No. Crawling permissions live in robots.txt, where rules can address specific crawlers, including Google-Extended for some of Google's AI products. llms.txt has no access rules at all.
What is the difference between llms.txt and llms-full.txt?
llms.txt is a short, linked reading list. llms-full.txt is a companion file some documentation platforms serve, such as Mintlify, which combines a whole documentation site into one file with the full text of every page. The llmstxt.org proposal defines llms.txt only.
Where should the file live and what format is it?
Usually at the root of your domain, at /llms.txt, served as plain text; the second version of the proposal also allows one at any path, covering the pages beneath it. The content is markdown: an H1 with the site name (the only required part), a blockquote summary, and H2 sections of links.
What should I put in my llms.txt?
Your most useful pages, grouped under a few H2 sections, each with a short note on what a reader will find there. For a software business that usually means product pages, pricing, documentation and a few core guides. Use plain language and leave out thin pages; the sitemap already covers completeness.
Is llms.txt an official web standard?
No. It is a proposal published at llmstxt.org and open for community input, with a public GitHub repository and a community Discord channel.
Will adding llms.txt improve my AI visibility?
Don't count on it for search: Google says AI text files aren't needed for its AI features. Measure the AI answers you care about directly, and treat llms.txt as a small addition for the agents that read it.
How do I generate one without writing it by hand?
Use the free llms.txt generator at /free-tools/llms-txt-generator. It reads your sitemap or homepage links, takes your main pages' titles and descriptions and drafts a file in the right format for you to review, edit and upload.

See if AI names you when customers ask who’s best.

Enter your website. In about two minutes, Rank.ai asks ChatGPT, Claude and Gemini 12 questions your customers ask and grades how often they name you.

  • Your grade out of 100How often AI names you, cites your site, and how high it ranks you.
  • Who gets namedEvery competitor in the answers, most named first.
  • The pages AI readsThe sources behind each answer.
  • Three fixesWhat to fix first, with a brief for the first page.