llms.txt & AI-readable content
llms.txt is a proposed Markdown file at a site's root that gives language models a short description of the site and a curated list of clean pages.
llms.txt is a plain Markdown file at a site's root that gives language models a short description of the site and a curated list of links to clean Markdown pages. Jeremy Howard of Answer.AI proposed it in September 2024. It is a community proposal, not a web standard, and Google says its Search does not use it and that creating it will neither help nor harm visibility.
- Origin
- Jeremy Howard (Answer.AI), 2024
- Level
- 401 · Expert
- Fits
- Startup, Small and mid-size, Scale-up
- Time to apply
- About two hours for a site of 50 to 200 pages
- What you need
- a list of the 10 to 30 pages a model should read first · a way to serve Markdown copies of those pages · access to the site root
llms.txt is a plain Markdown file, served at the root of a website, that tells a language model what the site is and which pages to read first. Jeremy Howard of Answer.AI proposed it on 3 September 2024. The proposal is maintained at llmstxt.org and in a public repository. It is a community convention. No standards body has adopted it.
What goes in the file
The file has a fixed order, and only the first element is required. The spec asks for an H1 with the site or project name. After that come an optional blockquote summary, optional paragraphs with details, and optional H2 sections holding lists of links, each with a short note. A section named Optional marks pages an agent can skip when it needs less context.

Here is an illustrative file for a clinic:
# Brightside Dental
> Family dental clinic in Leeds. Prices, treatments and booking rules are listed below.
Prices are in GBP and include VAT.
## Treatments
- [Implants](/implants.md): process, recovery, price range
- [Whitening](/whitening.md): options and safety notes
## Policies
- [Cancellation and deposits](/cancellation.md): rules and fees
## Optional
- [Team biographies](/team.md)
The proposal also asks sites to offer a clean Markdown copy of useful pages at the same address with .md added. Its page, last modified on 10 August 2026, describes discovery through rel="alternate" type="text/markdown" and rel="describedby" links, set either as HTML link elements or HTTP Link: headers. Real files go further. Anthropic’s developer docs file groups links under H2 and H3 headings such as Root URL, English and Messages, has no blockquote, ends every URL in .md and points to an llms-full.txt for the complete text. Stripe’s file also opens with instructions addressed to coding agents, for example to check the package registry for the latest version.
Size varies widely. Fetched on 9 October 2026, llmstxt.org’s own file is 637 bytes with 3 link lines, the Vercel file about 5 KB with 25, OpenAI’s about 6 KB with 40, Cloudflare’s 17 KB with 113, Anthropic’s 83 KB with 761 and Stripe’s 92 KB with 482. The large ones work as indexes for tools, not as pages a person reads.
llms-full.txt is not defined on the proposal page. Mintlify began generating both files for every docs site on 20 November 2024. Its post names Anthropic, Windsurf and Bolt.new as companies using /llms.txt, and ChatGPT and Perplexity as tools that can use it, which is Mintlify’s claim and not one from those companies. The full file puts the whole documentation set in one Markdown document, so a developer can hand a coding assistant everything in a single paste. The proposal page now lists Mintlify, GitBook, Yoast SEO, AIOSEO and Wix among the tools that generate the file. Yoast offers it free in Yoast SEO and publishes no data on its effect.
How it differs from robots.txt and sitemap.xml
robots.txt, specified in RFC 9309, says which URLs crawlers may request, and the RFC states that these rules are not access authorization. The RFC was published in September 2022, with four authors from Google. A sitemap, defined at sitemaps.org, lists the URLs on a site, with limits of 50,000 URLs and 50 MB per file and 2,048 characters per URL. llms.txt is a short list chosen by a person and written for a model to read.

| File | Question it answers | Status | Controls access? |
|---|---|---|---|
| robots.txt | Which URLs may a crawler request? | IETF standard, RFC 9309 | Asks crawlers to comply |
| sitemap.xml | Which URLs exist? | Shared protocol | No |
| llms.txt | Which pages matter most? | Community proposal | No |
Blocking AI crawlers is a robots.txt task. OpenAI separates GPTBot (training), OAI-SearchBot (ChatGPT search) and ChatGPT-User (user-initiated actions, where robots.txt rules may not apply). Anthropic separates ClaudeBot, Claude-User and Claude-SearchBot. Both document robots.txt control.
Do AI and search companies read it?
In the public statements reviewed for this page, none from Google, OpenAI or Anthropic says its product reads llms.txt to find, rank or cite sites. The record, in date order:
- 17 June 2025. Google’s John Mueller wrote on Bluesky that no AI system currently uses llms.txt, and said server logs show chatbots fetching pages but not the file (Search Engine Roundtable).
- 2025. Search Engine Journal reports that Gary Illyes and Amir Taboul said at a Search Central Live Deep Dive event in Asia Pacific that Google was not pursuing it. The report gives no date, and the account is secondhand.
- December 2025. An llms.txt file that Lidia Infante spotted on Google’s Search Central docs was removed within hours, per the same report. Dave Smart found similar files on developer.chrome.com and web.dev and suggested an automated CMS update, not a policy decision by the Search team.
- May 2026. Chrome’s Lighthouse 13.3 added an llms.txt audit in an experimental Agentic Browsing category that sits beside WebMCP checks. The page calls the file optional, a 404 is not a failure, and testing the category needs Chrome 150 or later.
- July 2026. Google’s guide to generative AI features, last updated 10 July 2026, says you do not need new machine-readable files. It adds that creating llms.txt is fine and will neither harm nor help visibility. A December 2025 version of Google’s AI features page likewise says no special files or markup are needed.
- 2 June 2026. Search Engine Journal reported that Mueller called its value “purely speculative for now” on Reddit (Search Engine Journal).
- 15 June 2026. Search Engine Journal reported that on Google’s Search Off the Record podcast he said a self-reported file cannot help a system tell sites apart, but could help an agent already on a site (Search Engine Journal).
OpenAI and Anthropic publish llms.txt for their own developer docs, and OpenAI’s crawler page links to its file. Neither company’s crawler documentation says its bots use the file when they visit other sites.
Measurement agrees. SE Ranking checked nearly 300,000 domains in a study published on 7 November 2025 and Search Engine Journal covered it on 20 November 2025. 10.13% of domains had the file, and it showed no measurable link to how often a domain was cited by AI answers in Spearman correlation, XGBoost and SHAP analyses. Removing the variable made the XGBoost model more accurate. The authors say the result depends on the dataset and models they used.
So llms.txt is not a ranking factor, and nobody has shown that it moves citations. Its clear use is narrower: developers and coding agents that load documentation on purpose.
What does make content readable to models
The evidence points to the pages, not the file. Vercel’s analysis of its network, published on 17 December 2024, found that GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and the other major AI crawlers it examined did not render JavaScript. In that month Vercel logged 569 million GPTBot fetches, 370 million from Claude, 314 million from AppleBot and 24.4 million from PerplexityBot, against 4.5 billion for Googlebot. The ChatGPT and Claude crawlers downloaded JavaScript files, 11.50% and 23.84% of their requests, but did not run them. Text that appears only after scripts run is invisible to the AI crawlers, while Google’s Gemini uses Googlebot’s rendering. Server-rendered HTML with real headings is the first fix. The technical SEO checks cover it.
Cloudflare measured its own page at about 16,180 tokens as HTML and 3,150 as Markdown, an 80% cut. A heading such as ## About Us costs about 3 tokens in Markdown and 12 to 15 in HTML. Since 12 February 2026 it converts pages to Markdown for sites on its Pro, Business and Enterprise plans that enable the feature, when a client sends an Accept: text/markdown header, and it adds an x-markdown-tokens header with the estimated token count. In the 2023 paper “Lost in the Middle”, Nelson Liu and colleagues found model accuracy is highest when relevant information sits at the start or end of the input. Put the answer first.
Google’s guide also says there is no requirement to break content into tiny pieces and no ideal page length, so write for readers first. The GEO paper by Pranjal Aggarwal and colleagues, accepted at KDD 2024, reports that changes to content raised visibility in generative engine responses by up to 40% on its benchmark. Authorship and sourcing, covered under E-E-A-T, support the same goal.
Teams that want an AI-readiness review alongside the rest of their content operation can start from the AI for real work practice.
How to apply llms.txt & AI-readable content, step by step
- Choose the audience. Decide whether the file serves developers and coding assistants that load documentation, or agents that visit your site for prices and policies. Result: a one-sentence scope that decides which pages belong.
- Pick the pages that answer real questions. List the pages a model needs for product, price, terms and contact questions. Skip archives and tag pages. Result: 10 to 30 links grouped under short headings.
- Make a clean Markdown copy of each page. Serve each listed page at its own address with .md added, without menus, banners or scripts. Result: every listed page can be read without parsing HTML.
- Write the file. Write one H1 with the site name, a blockquote summary of one or two sentences, then links under H2 headings with a few words on each. Put secondary pages under Optional. Result: a file a person reads in a minute.
- Publish it and check that it loads. Upload it so it resolves at /llms.txt with a 200 status, then open every linked address. Result: a working file that Lighthouse does not flag.
- Look for evidence of use. Search server logs for requests to /llms.txt and the .md pages by user agent, and compare them with the crawler names OpenAI and Anthropic publish. Result: a log-based answer to whether anything reads the file.
Examples
A developer documentation site
Illustrative. A payments API serves its 300 pages as Markdown and lists the 25 core ones under Quickstart, API reference and Webhooks, with changelogs under Optional. A developer pastes the file link into a coding assistant, which loads the listed pages. This is the use case the proposal was written for.
A dental clinic
Illustrative. A clinic lists treatments, price ranges, cancellation rules and booking steps with a Markdown copy of each. Google's guidance does not say this helps rankings, so the clinic treats it as a low-cost extra. The pages matter more: prices in plain text, the dentist's credentials and a visible review date.
A SaaS help centre
Illustrative. A software company with 400 help articles lists the 30 most visited under task headings and leaves out old release notes. Logs checked each quarter decide whether it stays.
When to use it
Use llms.txt when developers or agents load your documentation into a model, when your docs platform generates the file for free, or when you want one hand-kept map of your key pages.
When not to use it
Do not expect it to raise rankings or AI citations, and do not fund it as a project. The file is public, so list no private pages. On sites that change daily, generate it or drop it.
Common mistakes
- Treating it as a ranking factor. Google says the file neither helps nor harms visibility in its generative AI features.
- Listing every page. A sitemap with descriptions adds nothing a sitemap does not already tell a crawler.
- Linking to script-heavy HTML pages. The proposal asks for clean Markdown versions.
- Letting it go stale. Links to deleted pages or old prices feed wrong facts to anything that does read the file.
- Skipping the basics. Pages that load text only after JavaScript runs stay unreadable to the major AI crawlers Vercel examined.
FAQ
What is an llms.txt file?
An llms.txt file is a Markdown document at a site's root, usually /llms.txt, that names the site, summarises it and links to its most useful pages. Jeremy Howard proposed it on Answer.AI in September 2024 so models get a curated guide instead of raw HTML.
Does Google use llms.txt?
Google says Search does not. Its guide to generative AI features says creating the file is fine and will neither harm nor help visibility. John Mueller has said no AI system currently uses it. Chrome's Lighthouse has an optional llms.txt audit.
What is llms-full.txt?
A companion file that Mintlify began generating in November 2024. It puts all documentation text in one Markdown file so a coding assistant loads everything at once. The llmstxt.org proposal page does not define it, so formats vary between tools.
How is llms.txt different from robots.txt and sitemap.xml?
robots.txt tells crawlers which URLs they may request, and sitemap.xml lists the URLs that exist. llms.txt is a short, annotated list of the pages that matter, written for a model to read. It controls nothing and is not a standard, whereas robots.txt is specified in RFC 9309.
Do I need an llms.txt generator?
Not for a small site, where a text editor is enough. For large sites, Mintlify, GitBook and Yoast SEO generate and refresh the file.
Sources
- Jeremy Howard, The /llms.txt file (llmstxt.org)
- Answer.AI, /llms.txt: a proposal to provide information to help LLMs use websites
- AnswerDotAI, llms-txt repository
- Google Search Central, Guide to optimizing for generative AI features on Google Search
- Google Search Central, AI features and your website
- Chrome for Developers, Lighthouse agentic browsing: llms.txt audit
- Chrome for Developers, Lighthouse agentic browsing category
- Search Engine Roundtable, Google: No AI System Currently Uses LLMs.txt
- Search Engine Journal, Google says llms.txt is purely speculative for now
- Search Engine Journal, Google's Mueller says llms.txt can't help LLMs differentiate sites
- Search Engine Journal, Google's llms.txt guidance depends on which product you ask
- Search Engine Journal, llms.txt shows no clear effect on AI citations
- SE Ranking, LLMs.txt study of nearly 300,000 domains
- Mintlify, Simplifying docs for AI with /llms.txt
- Anthropic, Developer documentation llms.txt
- Stripe, Documentation llms.txt
- Yoast, llms.txt feature in Yoast SEO
- OpenAI, Overview of OpenAI crawlers
- Anthropic, Does Anthropic crawl data from the web
- IETF, RFC 9309: Robots Exclusion Protocol
- sitemaps.org, Sitemap protocol
- Cloudflare, Markdown for Agents
- Vercel, The rise of the AI crawler
- Liu et al., Lost in the Middle: How Language Models Use Long Contexts
- Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024)
Last updated Oct 9, 2026


