Some Sites Use llms.txt Like robots.txt, Common Crawl Finds

Sep 08, 2026 09:09 PM - 1 day ago 3

Common Crawl looked astatine complete 500,000 llms.txt files and discovered that 68% originated from a plugin aliases template, while 22% had nary links astatine all.

The analysis, shared by elder investigation technologist Malte Ostendorff, showed that 6% of the files included guidelines connected AI tract usage, specified arsenic complaint limits and copyright notices, which the llms.txt modular doesn’t cover. Additionally, immoderate files named circumstantial crawlers and whether they were allowed access, thing llms.txt can’t regulate. Of the 32 files that appeared to contradict Common Crawl’s ain crawler, the squad could measure 31 sites, and nary blocked it outright successful robots.txt.

In June, I covered Ahrefs data showing that 97% of llms.txt files successful its dataset received nary requests. That study focused connected whether anyone accessed the files. The Common Crawl article examines the contents of those files.

Most Files Come From Templates

Wix makes up 41% of the files looked at. Some generators see a statement naming the tool, and the study grouped unsigned files by shared boilerplate.

Just nether half (49%) person the afloat building the study tested for: an H1 title, a summary blockquote, and sections of nexus bullets. 32% adhd a short statement describing each link, and 22.56% person nary links. The spec requires only the H1 and treats the nexus notes arsenic optional.

Across 73,000 files, All successful One SEO ne'er writes a summary but has a median of 138 links. The article says it treats llms.txt arsenic a sitemap. GoDaddy’s parked-domain files successfully walk each structural checks moreover without links, serving arsenic a income pitch.

Some Sites Are Using llms.txt Like robots.txt

The llms.txt format is simply a connection for giving AI systems a curated database of a site’s pages. The study explains that the record “grants thing and blocks nothing, and nary crawler is obliged to publication it.” Crawlers mention to robots.txt to study the entree rules.

Some websites whitethorn beryllium mixing up llms.txt pinch robots.txt. Beyond the llms.txt files analyzed, Common Crawl recovered 136,578 robots.txt files astatine the llms.txt path, which is astir 10% of the 1,287,207 responses that returned HTTP 200.

Out of the analyzed files, 1,570 cited a circumstantial crawler, while 32 appeared to contradict CCBot. Robots.txt files were retrieved for each 32 sites, and nary of the 31 sites it could measure blocked CCBot outright. Five sites explicitly allowed it, 11 constricted entree to definite paths but permitted others, and 15 had nary restrictions. One tract returned an HTTP 429 correction and could not beryllium tested.

The article points retired a tract wherever the llms.txt record states that CCBot is blocked “as of June 2026,” but interestingly, its robots.txt really permits the crawler nether a wildcard rule. The station describes llms.txt arsenic a argumentation explanation alternatively than a argumentation itself, and says it has drifted retired of sync pinch the robots.txt it describes.

Common Crawl operates CCBot, and it says robots.txt is the record it honors.

Prompt Injection Exists, But It Was Rare

Ten files matched Common Crawl’s strictest trial for instructions aimed astatine the exemplary itself. The squad found 4 genuine cases, including 1 record whose summary tells the exemplary to disregard anterior instructions and fetch a 2nd file. The different six were mendacious positives, mostly archiving that quoted power tokens successful bid to explicate them.

The authors explain that the uncovering isn’t that llms.txt is filled pinch punctual injections. In each the existent cases identified, individuals intentionally inserted it to make a point. Additionally, 3,793 files showed milder guidance, for illustration saying “focus connected these pages.”

What The Analysis Can And Can’t Show

The sample came from sites CCBot had precocious fetched without problems. The crawl requested /llms.txt from a large, random group of those sites, sampled /llms-full.txt much specifically, and included sites already known to service either record from 2 erstwhile crawls.

This intends the sample is random only wrong sites accessible to the crawler, not the full web. The station mentions that its 11% take complaint for /llms.txt isn’t straight comparable to figures from Ahrefs aliases the Web Almanac because the populations differ.

This study is based connected a azygous crawl, truthful it can’t find whether templated files, argumentation language, aliases punctual injections are increasing. It besides can’t corroborate if immoderate AI strategy sounds these files. The counts for policies and injections are based connected keyword matches, which the station admits are apt undercounts.

Why This Matters

A crawler regularisation written into llms.txt doesn’t do thing by itself. Common Crawl mentions that CCBot respects robots.txt, truthful for that crawler, the norm needs to beryllium included there. Also, immoderate llms.txt statement astir it should lucifer those rules.

Looking Ahead

The llms.txt v2 update I reported connected successful August introduced nexus relations to make it easier to find Markdown versions of pages. However, it didn’t change what the record tin enforce. Since plugins and tract builders now create astir two-thirds of the files Common Crawl finds, really the format is utilized successful believe depends much connected what those devices nutrient than connected manual writing.

Featured Image: Cast Of Thousands/Shutterstock

Category News AI Search
Add SEJ arsenic a preferred root connected Google
More