The azygous largest AI crawler connected my website complete the past time was not an AI crawler. It arrived astir 1,500 times under Common Crawl’s name; it sent backmost nothing, and what it wanted was my SSH keys.
I went looking because of a number.
Cloudflare’s CFO Told Analysts Machine Traffic Could Reach 1,000 Times Human Traffic
Cloudflare’s Chief Financial Officer, Thomas Seifert, told analysts connected the company’s second-quarter net telephone that “if the existent trends continue, we deliberation successful 5 years, non-human postulation will beryllium arsenic overmuch arsenic 1,000 times arsenic overmuch arsenic quality traffic.” Then the statement that will seizure the headlines: “humans will beryllium a rounding correction connected the internet, not because quality postulation goes down, but that’s conscionable really accelerated we’re seeing non-human postulation grow.”
Two things worthy saying earlier anyone reaches for the pitchforks. First, Seifert added his ain caveat, unprompted: “with the large caveat that I person called it incorrect astatine each constituent on the way.” Cloudflare antecedently expected machine postulation to walk quality traffic successful 2027, and it happened successful May 2026. His errors person tally toward underestimating, which is the strongest statement for taking the projection seriously.
Second, the underlying measurement is real. Cloudflare’s ain station published the aforesaid week says less than half of each HTML page requests now travel from a human. I person nary statement pinch that. The instrumentality visitors are existent and they are the full taxable of this website.
The statement is astir what the number counts.
What One Day of Crawler Traffic connected My Own Website Looks Like
I pulled Cloudflare’s AI crawler position for nohacks.co for the 24 hours ending the evening of August 7. About 3,000 requests, of which astir a 3rd were unsuccessful, a fig up much than 1,000% connected the erstwhile period.
By crawler: CCBot 1,510. ChatGPT-User 375. ClaudeBot 296. Googlebot 245. PetalBot 107. Thirteen others sharing 353 betwixt them.
Image Credit: Slobodan ManicCCBot is Common Crawl’s crawler, the long-running non-profit web archive whose corpus trained a bully stock of the models everyone now argues about. On paper, it being my largest visitant is unremarkable.
Then I exported the paths.
It Asked for My SSH Keys, Not My Articles
Here are the most-requested paths successful that AI crawler traffic, pinch petition counts, precisely arsenic they came retired of the export:
- /.ssh/known_hosts (42 requests)
- /phpinfo.php (31 requests)
- /.boto (30 requests)
- /.env.production (29 requests)
- /.vscode/launch.json (28 requests)
- /.env.test (27 requests)
- /firebase-service-account.json (26 requests)
- /.gitconfig (24 requests)
- /server/.env (24 requests)
It continues for illustration that for a 100 paths: /id_rsa, /id_ecdsa, /private-key, /ssl/localhost.key, /key.json, /serviceAccountKey.json, /.aws/config, /actuator/configprops, /api/v1/env, /Dockerfile, /values.yaml, and /@fs/proc/self/environ, which is an effort astatine a known path-traversal bug successful a improvement server.
Across those 100 paths: 1,028 requests, 6.7 MB transferred, and zero referrals. The number of requests to thing I person really written rounds to nothing. The closest it came to my contented was /blog/wp-login.php, a WordPress login probe aimed astatine a website that has ne'er tally WordPress, and 2 requests for /blog/null.
That past item matters much than it looks. Whatever this is, it is not reference my pages earlier it asks for things. It is moving done a list, the aforesaid database it useful done everywhere, and my website is simply a statement successful a loop.
This is simply a credential scanner. Common Crawl follows links and fetches pages, and it has nary logic to inquire a podcast website for its Firebase work relationship key.
I could not verify the root addresses to beryllium impersonation, because per-request IP information is not thing I tin scope connected my plan. Common Crawl publishes the test: genuine CCBot postulation comes from documented reside blocks and reverse-resolves to hostnames ending successful crawl.commoncrawl.org. Someone pinch those logs tin settee it successful a minute. What I tin opportunity is what arrived, what it asked for, and really it was labelled: Cloudflare’s AI dashboard attributes this to Common Crawl arsenic the operator, and counts each petition toward my AI crawler totals.
Which leads to the portion that unsettles maine most. I went looking for these requests successful my information events and recovered thing astatine all, because the information log only records requests that travel a rule. I americium not blocking this traffic, truthful it passes through, gets served, and leaves nary mark. It appears successful precisely 1 spot connected my full dashboard: the AI crawler view, sitting successful the database beside ChatGPT-User and Googlebot, nether the sanction of a nonprofit investigation archive. A credential scanner is afloat legible to maine arsenic supplier postulation and wholly invisible arsenic a information event.
2 of Those Paths Are New, and They Are the Ones I Keep Thinking About
Buried successful that database are /.mcp.json, requested 30 times, and /.continue/config.json, requested 24.
Those 2 are supplier tooling configuration: an MCP server meaning and a coding assistant’s settings file. Both routinely clasp API keys and entree tokens, because that is what you put successful them to fto an supplier scope your services.
Someone has added supplier credentials to the modular secret-scanning wordlist. The aforesaid automated expanse that has been asking each website connected the net for /.env since astir everlastingly now besides asks for the record that lists which devices your agents tin telephone and what they authenticate with. Nobody announced that, and it happened fast. If you tally thing agentic, the wordlist arrived earlier astir group vanished penning their first MCP server.
Cloudflare Published the Correction Itself, the Same Week
The strongest counterweight to the earnings-call framing is successful Cloudflare’s ain engineering penning from the aforesaid week.
Their agentic-internet post says a batch of postulation from well-behaved bots is re-fetching pages that person not changed, and that this runs to billions of requests. In their words, “an tremendous magnitude of instrumentality effort, attached to nary result astatine all.”
Machine effort and instrumentality request are different quantities. My ain logs are a sharper type of the aforesaid constituent than I expected to find: the largest azygous contributor to my instrumentality postulation was not simply useless, it was hostile, and it still counted.
Meta crawling your website and ne'er sending thing backmost is the meaning of useless postulation if you are the personification who owns the website. I wrote astir that divided on August 1. A scanner wearing a investigation crawler’s sanction while it hunts for your unreality credentials is simply a class beneath that, and some onshore successful the aforesaid barroom connected the aforesaid chart.
So erstwhile the chart climbs, the mobility for a website proprietor is what the postulation really is.
Help Create the Problem, Market the Problem, Sell the Solution
It is clear what Cloudflare is positioning itself arsenic here, and it should beryllium called out. Help create the problem, marketplace the problem, waste the solutions. In the first week of August alone: a bot-traffic projection connected the net call, a blog station quantifying really overmuch of the web is nary longer human, an agent-readiness scanner to show you that you are not ready, an AI-visibility merchandise to people you, a span to expose your website’s devices to agents, and a default that starts blocking immoderate of those agents successful September unless you determine otherwise.
Every 1 of those products is simply a reasonable consequence to thing real. That is what makes the shape worthy noticing alternatively than dismissing. The institution measuring the problem, framing the problem, and trading the hole is 1 company, and they now ain some the metre and the valve.
I want to beryllium observant here, because I person backed a batch of what Cloudflare has done. Pay-per-crawl was the correct idea. Content Independence Day was the correct idea. Giving website owners a existent prime complete which machines get successful thumps a tribunal deciding it for them, which is what I based on erstwhile the Ninth Circuit took up that question connected August 4.
All of that tin beryllium existent astatine once. Cloudflare tin do immoderate bully things, immoderate directionally bully things, and immoderate things that look sketchy, astatine the aforesaid time. Most companies can. The correction is deciding they are the bully guys aliases the bad guys and past reference everything they do done it.
Go and Look astatine Your Own Logs
Take the postulation numbers earnestly and return the framing pinch the brackish it deserves. Machines are the mostly of requests. That is measured, and it is true.
Then open your ain crawler analytics and publication the paths, not the totals. Mine told maine 3 things I did not cognize this morning: that my largest AI crawler was a scanner, that it was burning megabytes of my bandwidth connected nothing, and that the wordlist it useful from now includes the config files it thinks my supplier tooling lives in.
None of that item is successful anybody’s projection. The measurement is. Fifteen 100 of these arrived astatine 1 mini website successful a azygous day, each 1 of them counting toward the thousand-to-one Seifert described to analysts, and not 1 of them wanted thing I wrote.
More Resources:
- Complete Crawler List For AI User-Agents [Dec 2025]
- US Publishers Demand Common Crawl Stop Scraping Their Content
- Should I Block AI Crawlers Or Measure Their Value First? – Ask An SEO
This station was primitively published connected No Hacks.
Featured Image: Lightspring/Shutterstock
English (US) ·
Indonesian (ID) ·