What Is llms.txt — and Does Anything Actually Read It?
llms.txt is a proposed file that tells AI systems which pages on your site matter. On the current evidence, almost nothing reads it. In May 2026 Ahrefs looked at 137,210 domains and found that 97% of the llms.txt files that existed received zero requests that month — not from AI bots, not from anything. Adoption is climbing fast and measured effect remains at zero. Publish one if you like; it costs minutes. Do not count it as an AI-visibility tactic, and do not let it displace the tactics that do have evidence behind them.
Our own measurement: on 2 September 2026 we ran the identical buyer-intent query — “best GEO / AI-visibility tools” — across Google, Perplexity and ChatGPT on the same day, and logged every domain each one cited. Between them they cited roughly 35 distinct domains. Exactly one,
promptwatch, appeared on all three. Whatever decides citations, it is not a shared, file-driven signal that every engine reads the same way. That result is why we treat single-file tactics like llms.txt with suspicion, and it is the backdrop to everything below.
Key Takeaways
- 97% of existing llms.txt files were never fetched at all in May 2026, across 137,210 domains (Ahrefs).
- AI bots never request the file on sites that don’t have one. They don’t go looking, which means the file cannot be a discovery mechanism.
- Adoption grew from 0.3% to 8.7% of the top 1,000 sites in twelve months (Rankability, June 2026) — roughly 29x growth in a signal with no measured payoff.
- Google’s John Mueller has publicly argued the format cannot help LLMs tell sites apart, because every site claims to be the best one.
- The AI labs publishing llms.txt for their own docs is evidence of publishing, not of reading. The two get conflated constantly.
What Is llms.txt, Exactly?
llms.txt is a convention proposed by Jeremy Howard in 2024, updated to a version 2 spec in August 2026 and documented at llmstxt.org. It is a plain markdown file at /llms.txt containing an H1 with your project name, a blockquote summary, and H2-delimited sections of curated links with short descriptions.
The reasoning behind it is sound. Web pages are built for humans, wrapped in navigation, cookie banners and ads, and an LLM extracting facts has to strip all of that away. A curated markdown index is genuinely easier for a machine to consume. Think of it as a hybrid of sitemap.xml (what exists) and robots.txt (what is permitted), aimed at language models rather than search crawlers.
The proposal is a good idea. The question this article is about is narrower and more awkward: is anything on the other end actually listening?
Does Anything Actually Read It?
This is where the evidence gets uncomfortable for the format.
In May 2026, Ahrefs analysed 137,210 domains using its own Web Analytics product, looking at real server-side request logs rather than surveys or guesses. Two findings matter:
97% of existing llms.txt files received zero requests that month. Not “few requests” — none. Nothing fetched them at all: no AI bot, no crawler, no human.
Zero requests arrived for llms.txt files on sites that don’t have one. AI bots never speculatively probe for the file. That single detail undercuts the whole premise, because a file nothing looks for cannot be a discovery mechanism. It can only be found by something already crawling your site, which by definition has already found you — and what actually gets you found is a different list entirely.
Of the 3% of files that did get fetched, AI bots accounted for 19.5% of requests, with GPTBot the largest single training crawler at 4.51% and an agentic indexer at 3.52%. Roughly 4% of traffic to those files was human — mostly SEO professionals and link-preview bots checking whether the file existed.
| What was measured | Finding | Source |
|---|---|---|
| Domains analysed | 137,210 | Ahrefs Web Analytics, May 2026 |
| Publishing an llms.txt | 28% of that sample | Ahrefs |
| Files that got zero requests | 97% | Ahrefs |
| AI-bot share of requests to the 3% fetched | 19.5% | Ahrefs |
| Requests for llms.txt on sites without one | Zero | Ahrefs |
One caveat the study itself invites: 28% adoption is far above every other published estimate, because the sample is sites using an SEO analytics product. That is a population unusually likely to have tried llms.txt. It makes the 97% figure more damning, not less — this is the group most motivated to make the file work.
What Do the Platforms Say?
No major AI platform has stated that its crawlers or models read llms.txt on third-party sites when deciding what to retrieve or cite.
Google has gone further than silence. On its Search Off the Record podcast, Search Relations’ John Mueller argued the format cannot do the differentiating job people expect of it, characterising the file as: “you’re telling these systems, like, I have the best website ever. And here are all of the pages that everyone must go to.” His point is structural rather than dismissive — self-reported importance is worthless as a ranking signal precisely because every site would claim it. He allowed one narrow exception: an agent already on your site might use the file to navigate.
The strongest argument in the format’s favour is on llmstxt.org itself, which notes that OpenAI, Anthropic and Google all publish llms.txt files for their own developer documentation, and that Chrome’s Lighthouse now audits for one.
That is real, and it is also the single most misread fact in this debate. Publishing an llms.txt for your own docs is not evidence that your crawler reads other people’s. A documentation team shipping a machine-readable index is making its own docs easier for agents to consume. It says nothing about whether the retrieval pipeline that decides citations consults the file on a site it is crawling. Those are different teams solving different problems, and the gap between them is where most llms.txt advice quietly goes wrong.
Adoption Questions
How many sites have one?
Depends entirely on the sample, and the spread is instructive.
Rankability checks the Tranco list monthly, requesting /llms.txt and /llms-full.txt and counting only verified text files returning HTTP 200 — excluding HTML pages, empty files and soft 404s. In June 2026 it found 8.7% of the top 1,000 (87 sites) and 5.6% of the top 10,000 (558 sites).
Independently, Casey Burridge’s HTTP Archive analysis put top-10,000 adoption at 5.61%. Two different methodologies landing within 0.01 points is about as good as corroboration gets in this field, and it makes the ~5.6% figure the one to quote.
Is adoption growing?
Sharply. Rankability’s top-1,000 figure went from 0.3% in June 2025 to 8.7% in June 2026 — close to 29x in a year.
That growth is worth sitting with, because it is growth in a signal with no demonstrated effect. Some of it is genuine bets on a future standard. Some is platform default: Shopify rolled the file out across its stores automatically, which moves the percentage without anyone deciding anything. And some is simply that it is easy to sell as a deliverable. Both counts are built on the Tranco list, which matters more than it sounds. Tranco exists because commercial top-site rankings turned out to be trivially manipulable: its authors demonstrated that a domain’s rank on Alexa could be changed through as little as a single HTTP request (Le Pochat et al., NDSS 2019). Adoption percentages measured against a manipulable list would be close to meaningless, so two independent counts agreeing on a hardened one is real corroboration rather than coincidence.
Does publishing one carry any risk?
Practically none. It is a small static text file. It does not affect indexing, it does not leak anything you have not already published, and it will not attract a penalty. It is not a crawler-control mechanism either — that job belongs to robots.txt, which is a published standard that crawlers demonstrably obey. The only real cost is opportunity cost — the hour spent on it, and the false sense that an AI-visibility box has been ticked.
What Has Actually Been Measured?
It is worth setting llms.txt against something that has been tested rather than proposed.
The clearest published attempt is GEO: Generative Engine Optimization by Aggarwal et al., accepted to KDD 2024. The authors built GEO-Bench, a benchmark of user queries across multiple domains, and measured what happened to a page’s visibility in generative-engine answers when the page itself was edited. Their headline result: the methods they tested boost visibility in generative engine responses by up to 40%. The edits doing the work were content-level ones — adding citations, adding quotations from credible sources, adding statistics.
That is the comparison worth holding onto. One approach has a benchmark, a published methodology and an effect size. The other has a proposal, a growth curve, and a 3% read rate. They are not really competing for the same afternoon, because only one of them has ever been shown to do anything.
For example: a page asserting “AI search is growing fast” and the same page rewritten to say “AI-search referrals grew X% between March and August 2026, per [named source]” are the same claim at different evidence densities — and that gap is precisely the class of change the GEO paper found effective. Writing a curated index at /llms.txt and hoping is not in that class at all.
What a decent llms.txt looks like
If you are going to spend the ten minutes, spend them on something an agent already crawling your site could use. The spec asks for four things, and most generated files supply none of them.
| Element the spec asks for | What a useful one contains | What most generators emit |
|---|---|---|
| Project name (H1) | The site’s name as an entity | The domain string |
| One-line summary (blockquote) | What the site is for, in a sentence | Omitted |
| Grouped sections (H2) | “Core”, “Guides”, “About” — a shape | One flat list |
| Links with descriptions | A sentence saying what each page establishes | Bare URLs, no descriptions |
Ours is live at citablepress.com/llms.txt if you want the whole thing rather than a summary of it. For example, its entry for the previous article does not just link the URL — it says what the page establishes (“what 400M citations show about which pages get quoted”), which is the only part an agent could not have worked out by crawling.
That is the actual test. A file that restates your sitemap in worse syntax and with no schema tells a machine nothing new. A file that says what each page establishes is the curated index the proposal describes — and it is still, on today’s evidence, being read by almost nothing.
A decision aid
| If you are… | What to do about llms.txt | Why |
|---|---|---|
| A solo publisher or small site | Hand-write one, once, then stop thinking about it | Ten-minute cost, no measured return |
| An agency or freelancer | Do not bill it as an AI-visibility deliverable | Nothing published supports the claim |
| A docs, API or app site | Worth more care than average | Mueller’s one conceded case: an agent already on your site |
| Choosing between llms.txt and page-level edits | Page-level edits, every time | Only those have a measured effect size |
| Being sold an “llms.txt optimisation” package | Ask for the read-rate data on their own clients | The 97% figure is the base rate they are arguing against |
So Should You Publish One?
Yes, if you understand what you are buying: a cheap option on a standard that might become meaningful, not a tactic with evidence behind it.
The honest framing is a wager with a tiny stake. Cost: minutes. Downside: none measurable. Upside: unproven, and currently indistinguishable from zero. That is a reasonable bet at that price — and an unreasonable one at any price above it.
What the evidence does not support:
- Treating llms.txt as an AI-visibility tactic in a strategy document
- Prioritising it over work with measured effect — ranking in conventional search, answer-first structure, cited statistics, unblocked crawlers
- Paying an agency for llms.txt generation as a line item
- Reporting it to a client as work that improves AI visibility
We publish one on this site. We also publish this article saying we expect nothing from it, and we will report it here if that ever changes — that is what our methodology commits us to.
The Bottom Line
llms.txt is a well-designed answer to a real problem that, on today’s evidence, nobody is asking for. Adoption is up nearly 29x; the read rate across 137,210 domains is 3%, and AI bots never go looking for the file at all. Google’s own Search Relations team has explained why the format cannot carry the weight being put on it.
Ship one if you want. Spend the afternoon you saved on something the evidence supports.