Google’s Mueller Says AI Crawlers Access Sitemaps & RSS In His Logs

Search Engine Journal by 4 min read 138x views
Google’s Mueller Says AI Crawlers Access Sitemaps & RSS In His Logs

Share Post

Google’s John Mueller says AI training crawlers normally don’t provision location owners a way to present a sitemap, and he suggests a default document name or an RSS nourish for sites that desire those crawlers to discover their content.

He made the comments alongside Martin Splitt in the Oct. 1 event of Google’s Search Off the Record podcast, titled “Do sitemaps motionless matter?” He additionally stated he has seen AI crawlers admission his sitemap and RSS records in his own server logs.

Private Sitemaps And AI Crawlers

Mueller stated location owners who desire to keep a sitemap personal can provision it an different document name, depart it out of robots.txt, and present it to Google directly. The downside is that another systems can’t discover it, he said, and Bing would likely need its own submission.

AI training crawlers are distinct since “they normally don’t have any benevolent of a Console or any setup anywhere you can present a sitemap file.” For sites that desire their satisfied in AI systems, he suggested “either you rod to the generic naming, call it sitemap.xml, or you concentration on RSS feeds.”

The sitemaps protocol says a robots.txt Sitemap row is autonomous of the user-agent line, so it isn’t tied to any one crawler’s rules. Feeds are easier to find, he noted, since they’re normally connected from a page’s HTML head.

“I’ve seen that happen in my server logs anywhere several AI crawler accesses my sitemap file,” he said, and he’s seen the identical alongside his RSS files. He didn’t name the crawlers and stated he doesn’t cognize whether AI companies document this or what they do alongside the files.

Llms.txt Isn’t A Sitemap Substitute

Asked whether llms.txt could substitute an XML sitemap, Mueller compared the Markdown document to an HTML sitemap and stated Google’s systems can’t use it as a sitemap since it lacks the strict format.

“I think the anticipation is bigger than the reality,” he said, allowing that hunt systems power peruse Markdown records someday. “Currently none of this happens.” He’s fine alongside sites trying it but wouldn’t depend on it.

In August, he stated the lone crawlers on his test sites that claimed to obtain Markdown were SEO tools. In May, I covered Google’s AI optimization guide listing llms.txt among the strategies sites don’t need for generative AI features.

Why A Valid Sitemap Can Show ‘Couldn’t Fetch’

Splitt closed by asking why Search Console reports “Couldn’t fetch” for a valid, community sitemap connected from robots.txt.

Mueller stated one logic is presenter load. Google’s systems may be too occupied to fetch the sitemap, and Search Console reports that as “Couldn’t fetch” too. The another is crawl demand, which can guide Google to skip the sitemap whenever its systems see no need to crawl much additional from a site.

“And the crawl petition is extremely frequently according to the perceived norm of a website,” he said, adding that “it’s not purely a specialized thing.” In February, he told a Reddit user that Google won’t use a sitemap if it isn’t convinced there’s new and crucial satisfied to index.

Search Console’s Sitemaps study assistance page lists low crawl petition among the reasons a sitemap can’t be fetched and says improved satisfied brings additional crawl demand. Other reasons it gives contain robots.txt blocks, unresolved manual actions, and incorrect URLs.

Why This Matters

A sitemap alongside an different name and no robots.txt admission stays hidden from systems it isn’t submitted to, according to Mueller, and AI training crawlers normally don’t recommendation a way to present one. He suggested a default sitemap name or an RSS nourish for sites that desire AI crawlers to discover their content. A robots.txt listing additionally tells crawlers that assistance the Sitemap directive anywhere the document is.

For a valid sitemap showing “Couldn’t fetch,” the two causes Mueller described, presenter burden and crawl demand, sit exterior the document itself.

Looking Ahead

Google’s Search Relations guide advised operating from what’s documented now fairly than preparedness about llms.txt. Server logs are anywhere he saw AI crawlers attain his files, and they’re anywhere you can inspect your own.

Featured Image: Tetiana Yurchenko/Shutterstock. Google wordmark: Source: Google.

Category News AI Search
Other Article Search Engine Journal
↑
Close Right Ads
Close Left Ads