On August 20, I published a leaf arguing that a website’s robots.txt and its genuine implementation are two distinct things, and I used my website as the example of getting it wrong. My robots.txt had spent months welcoming Bytespider by name, lengthy following I would have chosen otherwise, and I lone caught it during penning that page.
The next day, Cloudflare announced Bot Preference Sync, a merchandise that fixes exactly that.
Bot Preference Sync writes your robots.txt for you. Whatever bot guideline you set in Cloudflare’s dashboard gets turned into robots.txt entries and prepended to your file, inner # BEGIN Cloudflare Bot Preference Sync and # END Cloudflare Bot Preference Sync markers, alongside your existing satisfied preserved underneath. Cloudflare says it volition run from the liberated tier up, and that for new customers it volition be on by default.
As of September 13, there is no admission for Bot Preference Sync in Cloudflare’s bots changelog, which motionless ends at July 1, no citation of it anyplace in Cloudflare’s bots documentation, and no generated obstacle on my robots.txt. So what follows is a merchandise as announced.
Bot Preference Sync is good. My exception is that it hands a vendor the decision concerning what your website tells AI crawlers, in a form that cannot province what many websites really do, and it does that by default.
A Robots.txt That Contradicts Your Edge Is An Argument For Ignoring It
Cloudflare’s announcement of Bot Preference Sync says that “when your stated preferences and your enforced rules disagree, several crawlers treat it as a basis to disregard your preferences or try to bypass your enforced rules.”
Cloudflare does not say which crawlers, or how many, or how it knows. It sits in forefront of a ample portion of the web and sees the traffic, which makes it the one gathering in a stance to name them, and it names none. It appears akin item that could be true. A catalog would create it item I could check.
Cloudflare is confirming the disagreement I spent a entire citation leaf on in what courts say concerning blocking AI bots. A document that says one item during your border does another hands an disagreement to anyone who wants to disregard you.
That benevolent of gap is uncomplicated to create, since robots.txt and border implementation live in distinct places. The document is content you wrote once, likely a during ago. The implementation is a dashboard you changed at several item since. Nothing keeps them honest alongside all other, and nothing keeps either of them honest alongside what you would decide today. My document drifted from what I would decide for months, and I compose concerning this for a living.
I started fixing my robots.txt by hand on August 20, the day before Cloudflare announced Bot Preference Sync, and completed on August 26.
Bot Preference Sync Sets Policy Per Category, Not Per Crawler
Bot Preference Sync generates your robots.txt from three settings, which live under Security Settings and Configure AI bot policies: Search, Agent, and Training. Cloudflare’s records gives all three the identical options: obstacle on all pages, obstacle lone on pages alongside ads, or allow. The Bot Preference Sync notice describes Training differently, as a Disallow choice that writes a no-training row into the file. Cloudflare’s tracked bot catalog decides which crawlers autumn into which category, and the generated robots.txt entries prosecute from your three settings. You cannot exclude an idiosyncratic bot from the sync. Cloudflare’s stated answer for anyone who wants finer authority is to rotate the sync off and keep the document yourself.
My own crawler guideline does not fit any of Bot Preference Sync’s three categories. I authorize OpenAI’s GPTBot, Anthropic’s crawler and PerplexityBot. I obstacle Bytespider and meta-externalagent. Every one of those companies trains models. My regulation is a inquiry I ask per company: What am I getting in return? The archetypal three put my pages in forefront of group who ask assistants questions. The another two obtain and come back nothing. That is a endeavor decision made one crawler at a time.
Set Training to disallow and the document tells OpenAI not to train on satisfied I am blessed for OpenAI to train on. Set Training to authorize and nothing in the document separates Meta and ByteDance from anyone else, during my border returns 403 to both. There is no environment that describes what I really do. Cloudflare’s documented remedy is to toggle the sync off and go rear to penning the document by hand, which I had already been doing the day before Bot Preference Sync existed.
Cloudflare Published 4 Disclosure Conditions, And Attached Blocking To Them
Setting Training to disallow writes a no-training row into your robots.txt and blocks all AI crawler Cloudflare judges opaque. Cloudflare has published the conditions a crawler must encounter to evade that. For a bot that does the two hunt and training, current are the requirements in Cloudflare’s own words:
- It “must respect, via any mechanism, a ‘no training’ penchant in robots.txt”
- “They provision location owners a way to opt out of AI summaries”
- “They provision URL-level visibility into which pages were made accessible for training, as fine metrics on hunt results, so you can see how your satisfied was used for hunt and for training”
- “They can display publically that Disallowing Training does not hurt your traditional hunt results”
Miss those four conditions and Cloudflare treats the crawler as opaque, and opaque method blocked on all website that set Training to disallow.
No business is named anyplace in that list. Conditions two and four depict Google.
Condition four Google already meets. Its crawler documentation, final updated July 14 2026, says Google-Extended controls whether crawled satisfied trains Gemini models and evidence their answers, and that it “does not effect a site’s inclusion in Google Search nor is it used as a ranking indication in Google Search.” That is the community declaration circumstance four asks for, on Google’s own say-so.
Condition two is the one Google has no answer for. It asks for a way to opt out of AI summaries, and Google’s documentation on AI features says the controls are “nosnippet, data-nosnippet, max-snippet, or noindex,” all one of which limits what Search shows everywhere. The identical leaf sets the regulation that makes the two inseparable: to be shown as a supporting nexus in an AI Overview “a leaf must be indexed and eligible to be shown in Google Search alongside a snippet.” One toggle governs both. There is no environment that takes you out of AI Overviews and leaves your average hunt snippets alone, and Google’s crawler records does not citation AI Overviews anywhere.
Microsoft answered circumstance two in September 2023. Content tagged NOARCHIVE “will not be included in Bing Chat answers,” and specified satisfied “will motionless appear in our hunt results.” The satisfied stays out of the answer and stays in the index. The ask is sensible and it is uncomplicated to meet. Cloudflare’s July 2026 article on AI traffic options names BingBot alongside Googlebot as the mixed-use crawlers these conditions govern.
So Cloudflare, which sits in forefront of much of the web, has written downward what disclosure it expects from AI companies, and attached blocking to the answer. That seems fair to me. But the conditions were written by a vendor, the implementation is that vendor’s network, and the websites doing the blocking mostly clicked one toggle in a dashboard and never saw the conditions attached to it.
The accountability inquiry publishers have been asking concerning AI companies now points at Cloudflare too.
Cloudflare Says Bot Preference Sync Will Be On By Default For New Customers
On September 15, Cloudflare changes the defaults for all new domains: Training and Agent blocked on the pages that display ads, Search remaining allowed. Until afterward a new client who sets nothing gets no blocks at all, and Cloudflare says “the starting item volition not add any blocks on your behalf.”
Selecting “I monetize from pages alongside ads on this domain” during onboarding sets Training to Disallow for you. That is a inquiry concerning your endeavor model, and the answer to it becomes a published stance on AI training in your robots.txt.
On defaults mostly my stance is unchanged: Do not go for them without knowing what they do. The websites at hazard are the ones that never open robots.txt again and never peruse the stance written for them. They volition have a guideline on AI training that they did not write, cannot see, and could not have expressed in three settings anyway.
Every default gets accepted by group who never see it, Cloudflare’s included. Cloudflare already writes its analytics manuscript into free-plan websites unless you opt out, which its own blog announced in September 2025 and an August 2026 Hacker News thread rediscovered. The identical on-by-default form now reaches robots.txt, a document group open equal small frequently than a settings page.
Robots.txt Stops The Crawlers That Want To Be Stopped, And Nothing Else
A robots.txt regulation stops a crawler that chooses to be stopped and does nothing to one that does not. Bot Preference Sync volition create several group awareness protected.
The document is a request, and a petition lone plant on the well-behaved. I have measured the another benevolent on my own website: the biggest so-called AI crawler in my logs was hunting for credentials, asking for /.env and SSH keys under a nonprofit investigation archive’s name. No row in a content document was always going to inconvenience that.
So Bot Preference Sync is genuinely helpful for the honest fractional of the internet, which is a genuine half. The implementation is motionless the edge. The document is records of intent, which matters later, in a dispute, and not at the instant a petition arrives.
Compare Your Robots.txt To Your Cloudflare AI Bot Policies Before Bot Preference Sync Reaches You
Three checks obtain concerning 10 minutes and inform you whether Bot Preference Sync would alter what your website publishes.
Read your robots.txt, including the parts you wrote in 2023. Then open Security Settings, Configure AI bot policies, and compare. If those two disagree, you have the mismatch I had, and you get to decide who fixes it.
Then decide whether the three categories can province what you want. If your guideline is “open to everyone” or “closed to training,” they can, and Bot Preference Sync volition preserve you a job you should not have to do by hand. If your guideline is per company, they cannot, and the honest move is to rotate the sync off and keep the document yourself.
When Bot Preference Sync does attain your website, go appearance at what it wrote. A guideline document you have not peruse is a declaration person alternatively is making on your behalf.
More Resources:
- Should I Block AI Crawlers At Robots.txt Or Server Level? – Ask An SEO
- What Opting Out Of Google’s AI Search Features Means Now
- The Modern Guide To Robots.txt: How To Use It Avoiding The Pitfalls
This article was initially published on No Hacks.
Featured Image: PeopleImages/Shutterstock