ChatGPT’s page-fetching bot is disallowed by much sites than immoderate different AI bot of its kind. It besides reached disallowed pages connected much sites than immoderate different bot. OpenAI says robots.txt rules whitethorn not use to it because a personification asked for the page.
TollBit’s latest State of the Bots report has the numbers for the first half of 2026. Here’s what other the information shows astir really these crawlers behave and what it intends for your site.
Where The Bypasses Land
In the European sites discussed successful the report, astir 15% of identified AI page-fetchers reached URLs that the sites had marked arsenic disallowed.
This happens mostly pinch a fewer circumstantial agents. For example, ChatGPT-User, Bytespider, and Youbot each accessed disallowed pages connected astir half of the European sites that had explicitly listed them. Among these, ChatGPT-User reached the astir sites.
Sites Did Disallow It
Many of the newer page-fetching agents are hardly blocked astatine all. Only 9% of European websites disallow Claude-User, compared to 26% successful North America. Perplexity-User sits astatine 13% versus 26%.
Most of the newest agents person disallow rates successful the azygous digits crossed Europe, but ChatGPT-User stands retired arsenic an exception.
What OpenAI Says About The Rule
OpenAI’s crawler documentation says ChatGPT-User visits a page erstwhile a ChatGPT personification asks a question, and that because those actions are initiated by a user, robots.txt rules whitethorn not apply.
Perplexity says Perplexity-User mostly ignores the record for the aforesaid reason, but Anthropic has a different position and states that each 3 of its bots respect it, as we reported successful February. TollBit treats immoderate petition to a disallowed URL arsenic a bypass, sloppy of what the usability claims.
Why This Matters
A disallow statement for ChatGPT-User is simply a petition that OpenAI’s archiving says whitethorn not apply.
It’s important to look astatine a different facet here. According to OpenAI’s documentation, the supplier responsible for deciding if a tract shows up successful ChatGPT hunt results is called OAI-SearchBot, not ChatGPT-User. Sites that artifact some agents to forestall AI postulation person traded distant the visibility half of that woody and kept a fetching power that carries a carve-out.
Server logs aliases CDN records show what really arrived. The record only shows what you asked for.
Looking Ahead
Cloudflare is making immoderate updates to really it manages its crawler controls, moving the determination to the web layer. When it comes to the bots it recognizes, compliance is nary longer near up to the crawler itself. Starting from September 15, caller domains added to Cloudflare will person their Training and Agent crawlers blocked by default connected pages pinch ads, while Search crawlers stay allowed.
Whether the user-initiated loophole survives is the unfastened question. It rests connected the statement that requesting a page differs from a crawler taking it, and now each awesome assistants fetch pages this way.
Featured Image: Net Vector/Shutterstock
English (US) ·
Indonesian (ID) ·