Skip to main content
§01 · Blog / AEO

We removed llms.txt from our scoring engine. Here's why.

Ben LittleFounder, WhyIQPublished 3 August 202611 min read

llms.txt is a plain-text file that lists a site's most important pages in markdown so AI systems can read them without parsing HTML. If it is sitting on your AEO checklist, here is the short version: Google says its search ignores the file, and there is no published evidence it has ever earned anyone a single extra AI citation. Studies across 300,000-plus domains found no measurable link between publishing one and getting cited. We know this one from the inside, because our own scoring engine used to award points for llms.txt, and in June we deleted the bonus.

That deletion is the part worth reading about. A scoring product has an incentive to keep every check it ships: more checks feel like more value, and nobody churns over a bonus. We removed this one anyway, because the evidence said it was measuring effort, not effect. This post is the full data behind that call, plus the one place your llms.txt hour actually moves citations.

Comic panel: a proud robot butler holds up a glowing scroll labeled llms.txt while a stream of crawler robots rushes past him into a glowing library entrance labeled SEARCH INDEX.
The file is polite, well formatted, and addressed to robots that never stop to read it.

What Is llms.txt, Exactly?

A sitemap for robots that read markdown. That is the whole idea.

In September 2024, Jeremy Howard of Answer.AI proposed a simple convention: put a markdown file at yoursite.com/llms.txt that names your site, describes what it does, and links to your most important pages in a clean, LLM-friendly format. The pitch was reasonable. Language models have small context windows and HTML is noisy, so give them a curated table of contents. Sibling of robots.txt in spirit, aimed at AI readers instead of crawler permissions.

The idea spread fast because it was cheap, visible, and felt like doing something about AI search at a moment when everyone was anxious about AI search. Tooling vendors added generators. Agencies added it to deliverables. Checklists added it as a line item, usually near the top, because a file you can ship in ten minutes is the most completable task on any roadmap. By 2025 roughly 10 percent of major sites had one, and "llms txt" had become one of the most-searched terms in the entire AI-SEO category. It still is: about 5,400 searches a month as of August 2026, more than "answer engine optimization" and "generative engine optimization" combined.

5,400

monthly searches for llms.txt, more than AEO and GEO combined. Demand for the tactic outruns the evidence for it. Google Ads keyword data, August 2026

Does llms.txt Actually Work?

One in ten sites published the file. Then someone finally measured whether the robots read it.

Three independent lines of evidence landed in 2025, and they all point the same way. Adoption first: analyses put llms.txt on roughly 10.13 percent of surveyed domains, so this is not a fringe experiment that never got a fair shot. It got the fairest shot a convention can get, hundreds of thousands of live deployments across every industry.

Then the server logs. When researchers checked what AI crawlers actually request, llms.txt showed up in about 0.1 percent of crawler hits. The bots that were supposed to be the file's audience almost never open it. Otterly ran the experiment directly on its own properties and found no citation change. SE Ranking published the mechanical explanation of why it cannot help. And the largest analysis, covering more than 300,000 domains, found no measurable link between having the file and being cited by AI engines (Search Engine Journal, 2025). Adoption near 10 percent, effect indistinguishable from zero. In any other channel we would call that a failed A/B test and move on.

0.1%

of AI crawler requests touch llms.txt. The file's intended audience almost never opens it. Server-log analyses, 2025, n>300K domains

Comic split panel: on the left, a skyline of small websites each flying an llms.txt flag under the caption 10% adopted; on the right, a magnifying glass over a server log full of ordinary page requests with a single dim llms.txt line circled, under the caption 0.1% of hits.
One in ten sites flies the flag. The crawlers walk past almost every one of them.

Does Google Use llms.txt?

You rarely get a search engine to state a negative on the record. This time it did.

In June 2026 Google published an official guide to optimizing for its generative AI features, AI Overviews and AI Mode. It addresses machine-readable AI files by name, and the language is unambiguous: no special files, markup, or markdown are needed to appear in Google Search or its AI features, because "Google Search itself doesn't use them." The guide files llms.txt under exactly the category its fans hoped it would escape: an AEO hack that does not touch the ranking systems. John Mueller had already said the quiet part earlier: Google has no plans to support it.

Google even adds a courtesy clause: creating an llms.txt for other services is fine, it neither helps nor harms your Google visibility. Which is the polite version of the finding above. The file is not dangerous. It is inert. We covered the guide's full stop-doing list, chunking, AI-only rewrites, manufactured mentions, in Google just told you to stop buying AEO hacks. llms.txt is the list's headline act, because it is the one with a fan base.

The file is not dangerous. It is inert. And inert tasks that feel productive are the most expensive kind, because nobody audits them.

Comic panel: a stern librarian behind a towering desk labeled THE INDEX stamps NOT USED on a small llms.txt scroll while a tiny AI mascot watches nervously.
Google rarely states a negative on the record. For llms.txt, it did.

We Scored It. Then We Un-Scored It.

Our own scanner awarded llms.txt a bonus for months. Removing it was the most honest line of code we shipped last quarter.

WhyIQ's page scanner includes an AI Citability Index, a score that predicts how ready a page is to be cited from its on-page signals. Early versions gave a present llms.txt file a bonus inside the crawler-access dimension, nearly a fifth of that dimension's points. It felt defensible at the time. The convention was new, adoption was climbing, and scoring it rewarded people who kept up.

Then the evidence arrived, and it put us in a strange position: our published playbook was already telling readers the file had zero measured lift, while our scorer was still handing out points for it. Two parts of the same product disagreeing about the same fact. When Google's guide confirmed the studies in June, we deleted the bonus, and we went one step further: our recommendation engine is now explicitly instructed never to suggest creating an llms.txt file, because a model trained on two years of AI-SEO content will happily re-recommend it from habit. The habit is the hazard. The internet has generated far more words about llms.txt than evidence for it, and any system, human or machine, that learns from volume will keep repeating the tactic long after the data has retired it.

Comic panel: an engineer peels a glowing row reading LLMS.TXT +20 PTS off a scoreboard titled CRAWLER ACCESS, the row dissolving into cyan particles while the other checks stay lit.
The most honest line of code we shipped last quarter was a deletion.

Key takeaway

A score that rewards inert work is not neutral. It spends your optimization budget on the wrong line item. If your AI-visibility tool still awards points for llms.txt, ask what evidence the points are based on.

Why the File Can't Work (the Mechanics)

AI engines find pages through search indexes. A manifest file sitting next to your robots.txt never enters that pipeline.

The failure is structural, not a matter of adoption reaching critical mass. When ChatGPT, Perplexity, Claude, or Google AI answers a question, a retrieval layer queries a search index, pulls the most relevant live pages, and the model reads what came back. That is grounding, and we walked through the full mechanic in how AI decides what to cite. Notice what is absent from that pipeline: any step where the engine consults a site's self-published list of its own best pages. Retrieval selects documents by relevance to the query, ranked by the index. A manifest cannot inject a page into an index, cannot make it rank for a sub-query, and cannot vouch for its trustworthiness. Self-description is not a retrieval signal, for the same reason meta keywords stopped being one twenty years ago: it is the site grading its own homework.

There is also a quieter technical problem. The crawlers that matter fetch server-rendered HTML and largely skip JavaScript execution, and if they cannot fetch your actual pages cleanly, a tidy markdown summary of those pages solves nothing. The bottleneck was never format. It was access, relevance, and trust.

Should You Delete Your llms.txt File?

No. Keep it. Just stop counting it as work.

This is where the honest answer gets a little unsatisfying. The file is harmless, a handful of smaller AI tools do read it, and it costs nothing to keep. We publish one ourselves, the same way we publish a favicon. If yours exists, leave it alone. If it does not, ship one in the ten minutes it deserves and never think about it again.

The correction is to your budget, not your file system. llms.txt belongs in the hygiene column, next to the favicon and the 404 page, not in the growth column where agencies and checklists keep filing it. The test is simple: hygiene tasks are done once and never revisited, growth tasks earn a recurring line in your reporting. If your AEO retainer includes "maintain llms.txt" as a billed deliverable, you now know exactly what that line is worth.

What Should You Do Instead of llms.txt?

The same research that retired llms.txt also ranked what works. The top of the list is not glamorous.

The largest cross-study meta-analysis in the field, 54 studies scored by evidence strength (Zyppy, 2026), puts crawler and URL accessibility at 9.5 out of 10, the single strongest factor. The same analysis scores llms.txt at 2.0. So the first hour goes to a boring question: can the engines fetch your pages at all, as server-rendered HTML, without a JavaScript step they will not execute? We wrote up how often the answer is no in AI crawlers can't read your website.

The second hour goes to placement. Across a 548,000-page study, 44.2 percent of AI citation extractions came from the first 30 percent of body text (AirOps, 2026), so the sourced, self-contained answer belongs in your opening paragraphs, not in section four. The third goes to evidence density: adding real statistics and quotable claims lifted AI visibility by roughly 40 percent in the Princeton GEO study (KDD 2024). And the long game is reputational. Brand mentions across third-party sites correlate with AI citation at r=0.664, about three times the correlation of backlinks (Ahrefs, 2026). The full sequence, with the effect sizes on every move, is the AI Citability Playbook.

One more habit is worth stealing from this story: measure instead of assuming. We only caught our own llms.txt contradiction because we track what the engines actually cite, week after week, rather than trusting the checklist. Whatever tactics you run, answer engine optimization included, the loop only closes when something reads back the result.

9.5 / 10

evidence score for crawler and URL accessibility, the top-ranked citation factor. llms.txt scores 2.0 on the same scale. Zyppy 54-study meta-analysis, 2026

Comic panel: robot workers reinforce a glowing suspension bridge labeled CRAWLER ACCESS 9.5 carrying crawler robots toward a bright webpage, while a tiny pennant labeled llms.txt 2.0 sits forgotten on a shelf.
Build the bridge the crawlers actually cross. The pennant can stay on the shelf.

Frequently asked questions

What is llms.txt?

llms.txt is a plain-text markdown file served at a site's /llms.txt path that lists the site's most important pages and describes what the site does, so large language models can read a curated summary instead of parsing HTML. It was proposed by Jeremy Howard of Answer.AI in September 2024 as a community convention, similar in spirit to robots.txt but aimed at AI readers rather than crawler permissions.

Does Google use llms.txt?

No. Google's official guidance on its generative AI features states that Google Search ignores llms.txt and other special AI files, and that no machine-readable file is needed to appear in AI Overviews or AI Mode. Google's John Mueller has separately said there are no plans to support it. Publishing one neither helps nor harms your Google visibility.

Does llms.txt help with ChatGPT or Perplexity citations?

There is no published evidence that it does. Analyses covering more than 300,000 domains found no measurable link between having an llms.txt file and being cited by AI engines, and server-log studies show AI crawlers request the file in a fraction of a percent of their hits. ChatGPT, Perplexity, and Claude retrieve pages through search indexes, not through manifest files.

Should I delete my llms.txt file?

No. The file is harmless, costs nothing to keep, and a few smaller AI tools do read it. WhyIQ publishes one. The point is budgeting: publishing llms.txt is hygiene, like a favicon, not a citation lever. If it is on your AEO roadmap as a growth task, replace it with work that has measured effect, like rewriting your first paragraphs or earning third-party mentions.

What actually improves AI citation if llms.txt doesn't?

The evidence points at four levers. Crawler and URL accessibility is the top-scored factor across a 54-study meta-analysis: engines must be able to fetch your page as server-rendered HTML. Answer placement matters: 44.2 percent of citation extractions come from the first 30 percent of body text. Evidence density lifts visibility by roughly 40 percent (Princeton GEO). And brand mentions across third-party sites correlate with citation at r=0.664, roughly three times the correlation of backlinks.

Go deeper

For the mechanic behind every citation, see how AI decides what to cite. For the rest of Google's stop-doing list, see Google's own AEO myth-busts.

Find out which signals your page is actually missing.

WhyIQ scores the AI citability signals the evidence supports, crawler access, answer placement, statistical density, schema, freshness, and skips the ones it doesn't. Free first scan, results in about 2 minutes.

Get my free WhyIQ Score