Guides / llms.txt studies

Guide · every study, checked at source

llms.txt adoption: every study compared

Scans of the web's top sites find llms.txt on roughly 5% to 11% of domains, and no server-log or citation study has shown AI assistants making real use of it. The adoption figures disagree mostly because of the denominator (all domains or only the ones that answer), how each scan treats HTML pages served with a 200, files that platforms like Shopify create automatically, which top-sites list was used, and the date. Our own scan of the Tranco top 1,000 on Oct 9, 2026 found 10.2% of all domains, or 15.0% of the ones that answered.

People quote "10%", "5.6%", "28%" and "97%" for llms.txt as if they measured the same thing. They don't. Below is every llms.txt study we could find, read on its publisher's own page on Oct 9, 2026, sorted by what it measured: whether sites have the file, whether AI crawlers request it, and whether having it gets a site cited more. For what Google and the AI companies themselves say about the file, see Does llms.txt work?.

1. How many sites have an llms.txt file

Scans of top-sites lists first, newest first, then other samples. "Hit" is what each study counted as a real file. Figures as published, except where marked.

StudyHeadlineSites checkedWhat counted as a hitOut ofResult
AIVisibilityNerd (this site)
Oct 9, 2026
10.2% Tranco top 1,000 (list PY35J), crawled Oct 9, 2026 200, not HTML, not a soft-404, not a catch-all host, final URL still a .txt file Both: all 1,000 and the 679 that answered over HTTP 102 of 1,000 domains (10.2%); 15.0% of the 679 that answered. Another 120 returned 200 with an HTML page or empty body. The top 10,000 scan is running and will be added here.
Rankability
Sep 18, 2026
9.3% Tranco top 1,000 and top 10,000 (list V3YPN), crawled Sep 18, 2026 200, text, recognizable Markdown; HTML and empty replies rejected; catch-all hosts marked unknown All domains, including 474 of the top 1,000 it couldn't decide 9.3% of the top 1,000 (93) and 8.3% of the top 10,000 (830) serve llms.txt or llms-full.txt. Its June 2026 figures (8.7%, 5.6%) came from an older scanner and list; it says not to read them as a trend.
SEOmator
Sep 13, 2026
8.6% Cloudflare Radar top 100,000, crawled Sep 13, 2026 200 with a non-HTML, non-JSON body of 20+ bytes; pages that challenge bots rejected Both: all 100,000 and the 75,353 that answered 8,598 domains: 8.6% of all, 11.4% of those that answered. 10,775 more returned an HTML page with a 200. Shopify made 18.2% of the files found.
Mahiro Hirakawa (DEV Community)
Oct 6, 2026
15% Tranco top 1,000 (list 56WKN), crawled Oct 6, 2026 200 plus checks on the first byte, an H1 and the final URL Reachable only (742 of 1,000) 114 domains, 15% of the 742 that answered over HTTPS (11.4% of all 1,000, our arithmetic). 246 answered with a 200.
lintlab
Oct 2, 2026
6.6% Chrome UX Report top 1,000 origins, crawled Oct 2, 2026 200 with a plain-text or Markdown content type; body not inspected All 1,000 66 of 1,000 origins. 112 more returned a 200 with some other content type.
Casey Burridge
Jun 20, 2026
5.61% HTTP Archive crawl, Chrome UX Report rank buckets, crawled monthly, Jul 2025 to Jun 2026 HTTP Archive's llms_txt_validation field (real, parseable file; soft-404 HTML rejected) Sites HTTP Archive crawled in each bucket (7,504 in the top 10,000) 5.61% of the top 10,000 in June 2026 (421 of 7,504 crawled), up from 1.04% in July 2025. Top 1,000: 6.28%; top 1 million: 5.07%.
Chris Humphrey
Jun 18, 2026
7.4% Majestic Million top 10,000, crawled Jun 2026 Real Markdown file; HTML soft-404s rejected All 10,000 737 domains (7.4%). 1,050 returned a 200, but 313 of those were HTML pages. 12% of real files were plugin-generated.
Thunderbit
May 8, 2026
5.86% Tranco top 10,000 (downloaded May 6, 2026), crawled May 6, 2026 HTML, app shells and off-path redirects rejected All 10,000 (handling of unreachable domains not stated) 586 of the top 10,000 (5.86%); 75 of the top 1,000 (7.50%). It saw 1,606 replies with a 200; only 586 passed. llms-full.txt: 1.03%.
HTTP Archive Web Almanac 2025 (SEO chapter)
Jan 15, 2026
2.13% Every Chrome UX Report origin (about 15.4M mobile, 12.2M desktop), crawled Jul 2025 "Valid" file; rule not spelled out in the chapter All crawled sites 2.13% of desktop and 2.10% of mobile sites; 324,184 valid files on mobile. The whole web, a year earlier than the scans above.
SE Ranking
Nov 7, 2025
10.13% Nearly 300,000 domains; source of the list not stated Not stated All domains in its dataset 10.13% had an llms.txt file. Also tested citations (see below). Covered by Search Engine Journal on Nov 20, 2025.
Nicolas Sitter
Mar 21, 2026
6.3% 105,002 hotel websites in 7 countries, crawled Mar 2026 Not clear from the page Reachable hotel sites 6.3% of hotel sites; 12.4% in the US, 3.8% in France. WordPress SEO plugins made a third of the files.
Ahrefs
Jun 15, 2026
28% 137,210 sites using Ahrefs Web Analytics (its customers), crawled May 2026 200, Markdown not HTML, pages saying "404" or "Page not found" rejected All 137,210 38,360 valid files (28%). Ahrefs calls it an upper bound: its customers are SEO-minded.
Search Atlas
Jun 19, 2026
32.2% 18,621 active domains in its study cohort Clean file vs clean 404; 4,710 ambiguous replies dropped 13,911 domains left after dropping ambiguous replies 5,997 domains (32.2%). Search Atlas notes this is far above the ~10% in broad scans.
Originality.ai
Jul 2, 2026
36,120 Its own monitor of 3M+ sites (list not published), crawled Jun 2025 to May 2026 Not described Counts only, no percentage 36,120 sites with llms.txt by May 2026, up from 4,088 in June 2025 (8.8x).
Chris Green
May 11, 2025
105 Majestic Million (all 1M), crawled Feb and May 2025 Not described All 1 million 15 sites in February 2025, 105 in May 2025 (0.011%). The post also says "0.015%" for the 15 sites; 15 in a million is 0.0015%.

The scans of top-sites lists done in 2026 (Rankability, SEOmator, the DEV Community post, lintlab, Burridge, Humphrey, Thunderbit and ours) land between 5.1% and 11.4% of all domains checked. The two figures far above that, Ahrefs' 28% and Search Atlas' 32.2%, come from sites that use SEO software, which are much more likely than the average site to follow SEO advice. The Web Almanac's 2.1% covers the whole web in July 2025, when the file was newer.

2. Do AI crawlers ask for the file?

Server-log studies: who actually requested /llms.txt.

StudyHeadlineSites and windowResult
Ahrefs
Jun 15, 2026
97% unread 38,360 customer sites with a valid file; May 2026 97% of the sites with a valid file got no request for it in May. Bots made 96% of the requests that did arrive, and SEO audit tools outnumbered AI assistants and AI search bots.
EZY
Jul 27, 2026
7 reads 83 sites on its own platform; Apr 27 to Jul 19, 2026 OpenAI's crawlers fetched llms.txt 7 times and robots.txt 3,990 times. Anthropic 9 vs 3,120; Perplexity 0 vs 775; Googlebot 67 vs 5,125. Only Meta's crawler read it more than robots.txt (193 vs 172).
OtterlyAI
Feb 5, 2026
0.1% 1 site; 90 days 84 of 62,100+ AI bot visits asked for /llms.txt (OtterlyAI rounds it to 0.1%). The average content page got about three times as many AI bot visits.
Senthor
Jan 25, 2026
104 of 10M 10M+ identified AI requests it logged since Sep 2025 104 requests for llms.txt, none from an identifiable AI company.
MaxAEO (vendor)
Jun 11, 2026
41 requests 19 customer sites that publish the file; Feb to Apr 2026 GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended asked for llms.txt 41 times, on 3 of the 19 sites, against about 1.1 million AI page fetches.
Evil Martians
Jul 21, 2026
37 of ~770 Its own site; two months About 770 direct fetches of its llms files; 37 came from named AI assistants and crawlers. "The file is popular, just not with the audience it was built for."
DEJAN (Dan Petrovic)
Aug 15, 2026
4 fetches dejan.ai; 30 days llms.txt was fetched 157 times, but only 4 times by Google, Anthropic and OpenAI crawlers, which fetched robots.txt 1,150 times.
WISLR (Tony Castillo)
Mar 20, 2026
0 wislr.com; Feb 1 to Mar 20, 2026; 12,099 bot requests No AI bot requested /llms.txt. A follow-up on 3 sites through July (Jul 20, 2026) found CCBot, ClaudeBot and Googlebot fetching it, but not GPTBot. Follow-up
tools.belchamber.us
Sep 24, 2026
0 5 small sites run by the author; Sep 17 to 24, 2026 Named AI crawlers made 5,819 requests; none fetched llms.txt. Its 330 llms.txt requests were mostly the author's own monitoring.
AWRSHIFT
Jul 31, 2026
0 Its own site; 12 days in Jul 2026; 9,645 requests About a thousand AI crawler requests, no external fetch of llms.txt. All 39 fetches were its own tools.

The pattern holds from one site to Ahrefs' 38,000: the same AI crawlers that fetch robots.txt thousands of times rarely or never fetch llms.txt. Most requests for the file come from SEO audit tools, scanners like the ones in section 1, and site owners checking their own file. The newest data shows some crawlers starting to look (Anthropic's and Google's in WISLR's follow-up, Meta's in EZY's), while OpenAI's crawlers fetched it 7 times in 12 weeks across EZY's 83 sites and not at all in WISLR's follow-up.

3. Does having one get you cited more?

StudyHeadlineSampleResult
SE Ranking
Nov 7, 2025
No effect Nearly 300,000 domains No correlation between having llms.txt and how often a domain is cited by AI. Removing the llms.txt variable made its XGBoost model more accurate.
Search Atlas
Jun 19, 2026
No lift 13,911 domains; Copilot, Gemini, AI Mode, Grok, OpenAI, Perplexity No meaningful visibility advantage for sites with the file. Better-written files went with better visibility, but that mostly vanished once domain authority was controlled for.
MaxAEO (vendor)
Jun 11, 2026
+0.2 pt 240 matched pairs of domains Cited in 11.8% of tracked prompts with llms.txt vs 11.6% without, inside its own week-to-week noise.

Three studies, three methods (correlation and machine learning, a three-step control for domain authority, matched pairs), one answer: no measurable lift. All three are by companies that sell SEO or AI visibility software, and none is peer-reviewed.

Why the adoption numbers disagree

1. The denominator. A list of "top domains" is full of names that aren't websites: CDN hosts, ad servers, API endpoints. In our Tranco scan, 236 of the 1,000 domains didn't even resolve in DNS. Divide by all 1,000 and you get 10.2%; divide by the 679 that answered and it's 15.0%. SEOmator (8.6% vs 11.4%) and the DEV Community post (11.4% vs 15%) show the same gap. Rankability keeps every domain in the denominator, including 474 of its top 1,000 it couldn't decide either way; its 93 hits are 9.3% of all 1,000, but 17.7% of the 526 it could decide (our arithmetic).

2. Fake 200s. Many sites answer any unknown address with their normal HTML page and a "200 OK" status. Count every 200 as a file and adoption can more than double: SEOmator saw 20,171 domains answer 200 but only 8,598 serve a real file; Thunderbit 1,606 vs 586; Humphrey 1,050 vs 737; the DEV Community post 246 vs 114; we saw 120 HTML or empty 200s next to 102 real files. lintlab only checked the content type, not the body, so its 6.6% is not filtered the same way.

3. Files nobody wrote by hand. Platforms and plugins now create llms.txt for their users. Shopify alone made 18.2% of the files SEOmator found, and WordPress SEO plugins another 9.6%. Plugins made a third of the files in the hotel study and 12% in Humphrey's. That pushes adoption up for whichever list holds more shops and WordPress sites, and says nothing about whether anyone chose to publish the file.

4. The list. Tranco blends several rankings and includes infrastructure domains; the Majestic Million ranks domains by links; the Chrome UX Report (used by HTTP Archive and lintlab) counts sites real Chrome users visit; Cloudflare Radar ranks by DNS traffic. Ahrefs and Search Atlas measured their own customers. Each list holds a different mix of software companies (which adopt most: about one in five in our sample) and other sites.

5. The date. Adoption has grown fast. Burridge measured 1.04% of the top 10,000 in July 2025 and 5.61% in June 2026; Originality.ai's count rose 8.8x in the same year; Chris Green found 15 files in the whole Majestic Million in February 2025. A study from 2025 and one from late 2026 are measuring different webs. Rankability warns against reading its June and September 2026 figures as a trend, because its scanner and list changed in between.

The only earlier comparison we found, SEO Madman's (Aug 23, 2026), set five of these figures side by side and reached the same conclusion: "The gap is method, not disagreement about reality."

Our scan

On Oct 9, 2026 we requested /llms.txt from every domain in the Tranco top 1,000 (list PY35J), trying www. when the bare domain failed. A hit had to return a 200 with a non-HTML, non-empty body that wasn't an error page, a JSON error, or a catch-all reply (we also requested a made-up filename to catch hosts that answer everything), and the final address after redirects still had to be a .txt file. Result: 102 of 1,000 domains (10.2%), or 15.0% of the 679 that answered over HTTP. Among the domains we labelled by type, software and developer companies were far ahead of news sites, social networks and infrastructure hosts. One network and one run, so sites that block bots or our location count as misses. We're extending the scan to the top 10,000 and will add the result here.

What this means for you

If you want to publish an llms.txt, it costs nothing and does no harm. Just don't expect it to change how AI assistants see you: the log studies show the main AI crawlers barely ask for it, and the citation studies show no lift. Letting the crawlers in (robots.txt and firewall rules) and ranking in ordinary search matter more, as our llms.txt guide explains. If a tool you pay for scores you down for a missing llms.txt, compare what each tracker actually measures in our price index.

How we built this page

We searched for every published study of llms.txt adoption, requests or citations, then read each one on its publisher's own page on Oct 9, 2026, checking every figure here against the page's text. Dates are the publisher's. Not included: roundups that only repeat other studies' numbers, studies of hand-picked groups of a few hundred sites or fewer, and studies whose figures we couldn't read on the page. If we've missed a study, or a figure here has changed at its source, email hello@aivisibilitynerd.com.