What is llms.txt?
llms.txt is a Markdown file at a website’s root, at /llms.txt, that gives AI agents a short, curated map of the site: its name, a one paragraph summary, and links to the pages worth reading. Jeremy Howard of Answer.AI published the llms.txt proposal in September 2024 and revised it in August 2026.
It exists because web pages are built for people. Navigation, ads and scripts get in the way, and a model can hold only so much text at once, so the file hands an agent the short version in one known place. Agents mainly read it on demand, while working for a user, rather than as training data. It does not decide whether AI tools mention a company; that rests on what they can find and trust about it across the web, the wider work of AI search visibility.
What an llms.txt file looks like
It is ordinary Markdown in a fixed order:
- An H1 with the site’s name, the only required part.
- A blockquote with a short summary.
- Optional notes in plain Markdown, with no headings.
- H2 sections of links, each written as
[name](url): note. - An “Optional” section, by convention, for links an agent can skip.
An llms.txt example from a B2B site
This site’s own file, cut where marked:
# Katama
> Katama (listed on Google as Katama Consulting) is a B2B growth marketing agency in Boston that sells pipeline-as-a-service: an outsourced growth team that designs, builds and operates the engine that identifies a company's ideal buyers, creates demand, generates qualified conversations and converts that activity into measurable sales pipeline.
…
- Email: hello@katama.io
- Phone: +1 (774) 777-6454
- Team: Mark Zides (Founder and CEO), Hamza Hashim (Marketing Strategist), Katie Barnes (Client Partner), Alexis Feinberg (Social Media Expert)
…
## Company
- [How we work](https://katama.io/how-we-work/): the five-stage loop every engagement runs, the weekly cadence behind it, and the numbers it reports on
- [About](https://katama.io/about/): who Katama is, the four people who do the work, and what the agency believes
…
- [Home](https://katama.io/): the overview of the whole offer
It lists the same pages as the sitemap, one line each, and leaves out figures, prices, client names and testimonials, because this is the text a model is most likely to repeat as fact. It links to the ordinary HTML pages, which already carry their full text, rather than to Markdown copies.
What changed in version 2
The August 2026 revision mostly helps agents find the file, and clean copies of pages, without guessing:
- Files at any path. A file such as /docs/llms.txt covers the pages under it, and the most specific file applies.
- Markdown copies of pages can sit at page.html.md or page.md.
- Standard links to both. A page can point to its Markdown copy with
rel="alternate"and to its llms.txt withrel="describedby", in the page head or as an HTTP Link header set once at the server. - “Optional” is only a convention now, with no special meaning to software.
Who actually reads llms.txt?
Google Search: no
Google’s guide to optimizing for generative AI features lists llms.txt files among the things you can ignore, “as Google Search itself doesn’t use them”, and says keeping one “will neither harm nor help your site’s visibility or rankings in Google Search”. Googlebot may still fetch the file, which means nothing. Whether a page is quoted in AI Overviews depends on the page, because being cited in an AI answer is a different job from ranking.
Lighthouse: checks it, does not require it
Chrome’s Lighthouse flags a page only “if a server error occurs when attempting to retrieve the llms.txt file”. A missing file is marked not applicable, “as providing the file is optional at the moment.”
ChatGPT, Claude and Perplexity: not documented
OpenAI’s page on its crawlers, Anthropic’s help center and Perplexity’s documentation all point site owners to robots.txt, and none says those crawlers read llms.txt on other sites. Each company publishes an llms.txt for its own developer documentation, which is a different thing.
Coding agents: yes
The proposal says the files “are used most heavily for software documentation, where coding agents follow them to find API references and tutorials.” Documentation platforms such as Mintlify and GitBook generate one for the sites they host.
llms.txt vs robots.txt and sitemap.xml
| robots.txt | sitemap.xml | llms.txt | |
|---|---|---|---|
| What it does | Says which crawlers may fetch which paths | Lists every page for search engines | Guides AI agents to the pages that matter |
| Who reads it | Search engine and AI crawlers | Search engines | AI agents and tools that choose to |
| Can it block crawlers | Yes, for crawlers that respect it | No | No |
| Used by Google Search | Yes | Yes | No |
llms.txt has no allow or disallow rules, so a Disallow line in it does nothing. To keep content out of AI training, block training crawlers such as GPTBot and ClaudeBot in robots.txt.
Where AGENTS.md, MCP and schema fit
- AGENTS.md gives a coding agent instructions for working inside a software project, not a website, and which instruction file an agent loads depends on the tool.
- MCP, the Model Context Protocol, connects AI applications to outside tools and data. An agent calls an MCP server; it reads an llms.txt.
- Schema.org markup describes a page from inside it, as LocalBusiness markup does for a local firm. llms.txt describes the whole site from outside its pages.
Does a B2B company need one?
| Your site | Worth it? |
|---|---|
| A product with an API or developer documentation | Yes, with Markdown copies of the pages it links to |
| Software with a help center customers rely on | Likely, updated with each release |
| A services firm or agency | Cheap, with little to expect: keep it short and accurate |
| Anyone hoping to rank higher in Google | No |
The risk is not a penalty. It is a stale file, or a claim you would not want repeated, read back to a buyer by an assistant.
How to create an llms.txt file
- Pick the pages that answer what you do, for whom and how to reach you, plus any documentation or policies. Leave out redirected, gated and noindexed pages.
- Write the summary as the paragraph you would want an assistant to repeat about you: plain facts, nothing that expires this quarter. Give each link a short note on what the page holds.
- Publish it as UTF-8 plain text at the root, or at the path it covers, and check that it loads without a redirect, a login or a bot challenge.
- Point to it from a comment in robots.txt, as this site does, and from a
rel="describedby"link or header. - Keep it in step with the sitemap by changing both in the same edit. Here that rule sits in the instructions a coding agent follows on every edit to the site.
- Test it as the proposal suggests: give an assistant only your llms.txt and ask it questions about your business.
How to see whether anything reads it
Lighthouse cannot tell you; your server or CDN logs can. Filter the requests for /llms.txt by user agent. ChatGPT-User, Claude-User and Perplexity-User fetch because a person asked an assistant something, while GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot crawl on their own schedule. A request means the file was fetched, not that you were cited.
● About the author
Hamza Hashim is the marketing strategist at Katama. He owns the channel mix and the order it gets built in: what ships first, what waits for evidence, and what gets stopped when the evidence says stop. Meet the team.