llms.txt File: What It Is, How to Build One, and Whether It Actually Works
By Ghost Writr · · 11 min read
If you’ve added an llms.txt file to your site hoping it would get you cited in ChatGPT or AI Overviews, here’s the direct answer: there’s no verified evidence it does that, and the man who leads Google’s search relations team says AI services don’t even check for it. That doesn’t mean the file is pointless — it has a real, narrow purpose — but it’s not an SEO lever, and treating it like one will waste your time.
This article explains exactly what llms.txt is, how the spec is structured, how it differs from robots.txt and sitemap.xml, how to build one if you decide it’s worth doing, and what the actual evidence says about whether it moves anything.
What is an llms.txt file?
llms.txt is a proposed root-level Markdown file — placed at a URL like https://www.your-website.com/llms.txt — that provides a curated index of a site’s most important pages, including their URLs, titles, and short descriptions. Think of it as a hand-picked table of contents for a website, written specifically so a language model can scan it quickly and understand what the site is about and where its most important content lives.
It’s a young idea. The format was proposed by Jeremy Howard, co-founder of Answer.AI, in September 2024. That’s important context: this is not a web standard with years of adoption behind it, and it’s not something a standards body has ratified. It’s a proposal that some tools and sites have chosen to support.
Crucially, llms.txt is designed primarily for inference time — the moment an AI agent or chatbot is actively trying to answer a user’s question — rather than for training a model or for broad web indexing. That distinction matters more than most explanations of llms.txt let on. A file aimed at inference time only helps if an AI system actually fetches it live, at the moment someone asks a question that touches your site. It does nothing for how a model was trained months or years earlier, and it does nothing to get your pages crawled and indexed the way a sitemap does. If you’re hoping llms.txt will somehow retroactively teach an AI model about your business, that’s not what it’s built to do.
The file structure: what the spec actually requires
The official llms.txt specification defines a strict Markdown structure, and getting this right matters if you want the file to be machine-parseable rather than just a text dump. The structure is:
- H1 title (required): The name of the site or project, as a single top-level Markdown heading.
- Blockquote summary (optional): A short, one- or two-line description of what the site does, formatted as a Markdown blockquote.
- Explanatory paragraphs (optional): Additional context in plain prose, if the summary alone doesn’t cover what an AI system needs to know.
- H2 sections: Headings that group related links by theme — for example, “Documentation,” “Pricing,” or “Guides.”
- Bulleted Markdown links with descriptions: Under each H2, a list of links in standard Markdown link format, each followed by a short description of what that page contains.
Here’s a hypothetical example of what a minimal, spec-compliant file might look like for a small software company:
# Acme Analytics
> Acme Analytics is a dashboard tool for tracking website traffic and conversions.
## Documentation
- [Getting Started](https://acme.example/docs/start): Setup guide for new accounts.
- [API Reference](https://acme.example/docs/api): Full endpoint documentation.
## Optional
- [Blog](https://acme.example/blog): Product updates and industry commentary.
Note the “Optional” section name — the spec allows an H2 called “Optional” specifically for links that are useful but not essential, so an AI system with limited context space knows which links it can skip.
llms-full.txt: the companion file
The specification also defines a longer companion file called llms-full.txt, which bundles a site’s actual page content — not just links and descriptions — into one large Markdown file sized to fit within an AI model’s context window. Some documentation platforms, including Mintlify, support or integrate llms.txt as part of how they present docs to AI systems.
How llms.txt differs from robots.txt and sitemap.xml
This is where a lot of confusion happens, and it’s worth being precise. robots.txt tells crawlers what they can and cannot access on a site, and well-behaved bots enforce those rules by staying away from disallowed paths. sitemap.xml lists all of a site’s pages so search engines can discover and index them systematically. llms.txt does neither of those things — it merely suggests a curated shortlist of pages for AI models to consider, and no crawler or search engine is obligated to honor it.
That non-enforcement point is the crux of it. A robots.txt file controls what crawlers are allowed to access; llms.txt cannot stop an AI system from using your site’s content, and it cannot force any AI system to use it. It’s a suggestion, not a rule.
| File | Purpose | Enforced? |
|---|---|---|
| robots.txt | Controls crawler access to pages | Yes, by well-behaved bots |
| sitemap.xml | Lists pages for search engine discovery/indexing | Followed by search engines as a discovery aid |
| llms.txt | Suggests a curated shortlist of pages for AI models | No — non-binding |
It’s also worth stating plainly what llms.txt does not do: it does not affect Google rankings or traditional SEO, and it does not block AI systems from accessing other parts of your site — that access-control role still belongs entirely to robots.txt.
Why the idea exists: parsing content for a context window
The underlying motivation makes sense even if the execution is unproven. A typical webpage is full of navigation menus, JavaScript-rendered widgets, cookie banners, and layout markup that a large language model has to wade through to find the actual content. llms.txt is meant to strip that noise away, handing a model a clean, pre-summarized map of a site’s key pages so it can parse content faster and prioritize what matters, rather than reconstructing meaning from HTML clutter. That’s a reasonable engineering idea. Whether AI systems are actually using it that way in practice is a separate question — and the evidence on that is not encouraging.
Does it help with AI search visibility (GEO/AEO)?
Generative engine optimization and answer engine optimization are about improving how often your content gets cited or surfaced in AI-driven search — tools like AI Overviews or conversational assistants. Given llms.txt’s stated design intent (helping an AI parse and prioritize content at inference time), it’s easy to see why site owners chasing GEO/AEO gains have gravitated toward it. But intent and effect are different things, and the honest answer is that the effect has not been demonstrated.
Does it actually work? The evidence
This is the part most explainers skip, and it’s the part that matters most.
Google’s John Mueller has stated on Bluesky that none of the AI services have said they use llms.txt, and that server logs show they don’t even check for it — he compared it directly to the abandoned keywords meta tag, a once-popular SEO element that search engines eventually stopped reading entirely. That’s about as direct a rebuttal as you’ll get from someone at Google’s search relations team.
Field data backs this up. WolfPack Advising reports that across more than 100 sites where they added llms.txt, they have never seen it move AI citations on its own. And a study of 137,000 sites found that 97% of llms.txt files were never even requested by a bot. Combined with Mueller’s server-log comment, that’s a consistent picture from two independent angles: AI systems generally aren’t fetching the file, and where site owners have installed it, they aren’t seeing measurable citation gains.
None of this proves llms.txt will never matter — adoption stories change, and this is a young, informally governed proposal rather than an established standard. llms.txt is not accredited by any official standards body such as the W3C, and its actual usage by AI systems at inference time remains unclear, with limited analytics and attribution data currently available to measure impact one way or the other. That’s an honest state of uncertainty, not a green light or a dismissal.
Who’s actually adopting it
Despite the thin efficacy evidence, some real adoption exists, and much of it shows up in developer tooling. Major platforms including Mintlify, Cloudflare, and Cursor have integrated llms.txt support. That pattern makes sense: documentation sites and coding tools are exactly the use case llms.txt was designed for — a coding agent working inside a tool like Cursor benefits from a clean, curated map of a docs site far more than a general AI Overview snippet would. If you run a documentation site for a developer product, publishing an llms.txt file is a reasonable, low-cost way to make that content easier for coding agents to consume — even without evidence it improves general AI search citations.
How to create an llms.txt file
Building one doesn’t require special software, though tools exist to speed it up.
- List your priority pages. Decide which pages best represent what your site or product does — usually docs, key product pages, pricing, and top guides.
- Write the file in plain Markdown. Open a plain text editor and follow the spec: an H1 with your site name, an optional blockquote summary, H2 sections grouping links by theme, and bulleted Markdown links with short descriptions.
- Use a generator if you’d rather not write it by hand. A Vitepress plugin exists for documentation sites built on that framework, and various commercial tools can generate an llms.txt file automatically from a supplied URL. Some CMS and documentation platforms now offer llms.txt generation alongside their existing sitemap.xml output.
- Upload it to your root directory. The file needs to live at the root of your domain —
/llms.txt— the same level as robots.txt and sitemap.xml, not nested in a subfolder. - Add an llms-full.txt if you want to go further. For sites with substantial documentation, a companion file bundling full page content into one context-window-sized Markdown document follows the same logic as the index file, just deeper.
- Track it, but track it honestly. If you want any read on whether AI-referred visitors are landing on pages linked from your llms.txt file, use UTM parameters on those links and watch for AI-attributed traffic in your analytics — though be aware this measures traffic from citation links, not confirmation that a model actually parsed the file itself.
None of this is difficult. The honest question isn’t “can I build one” — it’s “should I spend the afternoon on it instead of something with proven return.”
Should you bother?
If you run a documentation site, a developer tool, or anything coding agents like Cursor might read directly, llms.txt costs little to add and fits a use case with real, if modest, adoption. If you’re a general business hoping it’ll boost citations in ChatGPT or AI Overviews, the current evidence doesn’t support that expectation — Google’s own search advocate says AI services aren’t checking for the file, and WolfPack Advising’s own report, covering more than 100 sites where it added llms.txt, found no citation movement from doing so. Build one if the developer-docs use case applies to you. Otherwise, put your time into content and internal linking that AI systems can actually find through the paths they’re proven to use.
FAQ
Is llms.txt an official web standard? No. It’s a proposal, first put forward by Jeremy Howard of Answer.AI in September 2024, and it is not accredited by any standards body such as the W3C.
Does adding llms.txt hurt anything if I add it anyway? There’s no evidence it damages SEO or AI visibility — it simply sits at your root directory as an unenforced suggestion. The cost is your time, not your rankings.
Will llms.txt stop AI from scraping my content? No. It has no access-control function at all — that’s robots.txt’s job.
Do ChatGPT, Google, or Perplexity officially say they use it? No AI service has stated that it uses llms.txt, according to Google’s John Mueller, who noted server logs show AI services don’t even check for the file.
What’s the difference between llms.txt and llms-full.txt? llms.txt is a short curated index of links with descriptions; llms-full.txt is a much larger file containing the actual page content, sized to fit an AI model’s context window.