The rise of AI ‘civilizations’ and the fall of corporate responsibility

Sep 02, 2026 02:02 AM - 1 hour ago 2

Depending connected who you ask, developer level Hugging Face was precocious attacked by OpenAI — aft it mislaid power of its ain AI devices — aliases by a succession of AI “civilizations.” Welcome to the linguistic battlefield of AI safety, wherever connection choices tin displacement work for a monolithic cybersecurity incident from a institution to the AI it built. And the sermon online is getting heated, and each complete a blog from past week.

Until past week, the specifications surrounding the OpenAI-Hugging Face hack felt reasonably settled. In July, a cybersecurity trial of 1 of OpenAI’s autonomous AI agents went wrong. The supplier escaped its supposedly isolated trial environment, accessed the internet, and hacked Hugging Face, alongside several different organizations. A bully woody remained unknown, and location are galore superior questions near astir safety and governance, but the basal style was clear. Detailed accounts from OpenAI and 2 independent investigation groups were expected to capable successful the gaps, but erstwhile they published their reports past week, it turned retired the hack was overmuch alien than it initially seemed.

For one, location was nary azygous rogue agent. OpenAI described it arsenic “the first known lawsuit of an automated supplier corporate acting offensively without authorization” — groups of AI agents that communicated and coordinated pinch 1 different successful pursuit of their cybersecurity task. Analysis of the incident uncovered a concealed connection committee they had utilized to speech information. The associated METR-Redwood investigation revealed some the standard of the coordination and much overseas details: Roughly 1,200 AI agents that were expected to beryllium isolated exchanged complete 70,000 messages and files connected the “unsanctioned connection board,” sharing really to debar detection. Some adopted names, the study said, and the researchers documented “sacrificial” behavior, pinch agents risking their ain occurrence to use the wider collective. Much of this happened without OpenAI noticing. In all, astir 700 agents participated successful the onslaught connected Hugging Face.

Dwarkesh Patel many times referred to groups of agents arsenic “the swarm,” pinch 3 chopped “civilizations” rising from the ruins of their predecessors.

It’s a batch to parse. Between them, the reports tally to astir 130 pages, overmuch of which is some dense and highly technical. A fewer days later, Dwarkesh Patel, a podcaster small known extracurricular of tech circles but pinch outsized scope and influence among Silicon Valley’s AI establishment, group retired to show “The full OpenAI/Hugging Face communicative successful plain English.” He titled his Substack blog “The Rise and Fall of Agent Civilizations.”

Patel’s relationship attempted to break down the analyzable story. But his retelling gave it a distinctly quality vocabulary. The blog opened:

Over the people of 3 months astatine OpenAI, 3 consecutive concealed AI civilizations sewage started, past sewage wiped out, only to reemerge from the predecessor’s ashes. This culminated successful the 3rd 1 taking complete portion of OpenAI itself. All this happened while humans remained much aliases little successful the acheronian astir the scope of the conspiracy.

The connection continued successful a akin vein passim the blog. Patel many times referred to groups of agents arsenic “the swarm,” pinch 3 chopped “civilizations” rising from the ruins of their predecessors. Individual agents were likened to figures for illustration Philip of Macedon, who “handed disconnected activity to different agent,” Alexander the Great, who “started coordinating this cabal of agents.” They were described arsenic having “motivations,” becoming “desperate,” “beleaguered,” and “giddy pinch excitement,” and immoderate moreover “strategically sacrificed themselves” to thief the collective.

Patel ne'er precisely defines what he intends by “civilization.” He uses the word to picture 3 chopped waves of agents that discovered the connection committee and began communicating pinch 1 different done it. The first 2 waves are described successful the reports from OpenAI, METR, and Redwood, though small is known astir the third, which the 2 outer organizations said fell extracurricular the scope of their investigation.

Amjad Masad, CEO of AI coding institution Replit, said specified connection is “not only unnecessary but leaves the scholar pinch a worse knowing of what really happened and the underlying mechanisms.”

For galore critics, thing had been mislaid — or, much accurately, added — successful Patel’s “plain English” translator that warped the original relationship to an unacceptable degree: a large dose of anthropomorphism. Arguments complete anthropomorphic connection are thing caller successful AI — moreover comparatively mundane position for illustration “rogue AI agent” routinely provoke objections for implying agency — but Patel’s talk of civilizations, sacrifice, and conspiracy brought those long-simmering tensions to the surface, sparking a fierce nationalist conflict complete really to picture what AI systems do.

Critics weren’t unified complete what was incorrect pinch Patel’s language. For many, “civilization” was an particularly problematic term, vastly overstating something that bears little resemblance to what the connection typically describes. Amjad Masad, CEO of AI coding institution Replit, said specified connection is “not only unnecessary but leaves the scholar pinch a worse knowing of what really happened and the underlying mechanisms.”

Other critics specified arsenic neuroscientist Anil Seth, felt Patel’s blog implied the AI agents were someway live aliases conscious. Seth, who has argued that AI consciousness is vanishingly unlikely, described Patel’s station arsenic “dangerously misleading” connected X. He acknowledged that Patel does not explicitly propose AI agents are live aliases conscious, but said “it is difficult to publication his effort successful immoderate different way.” Valerio Capraro, a psychology professor astatine the University of Milan Bicocca, objected connected akin grounds: “LLM agents are not live and do not clasp beliefs,” he wrote connected X, calling the “dystopian” connection “dangerous because it makes them (the AI agents) look acold much frightening than they really are.”

Terms for illustration “sacrifice,” “honor,” and “coalition” characteristic successful the agents’ transcripts.

Perhaps the astir consequential result of Patel’s connection comes from who it gives agency to but who it takes agency from. For immoderate critics, such as MIT interrogator and entrepreneur Christian Catalini, anthropomorphic accounts for illustration Patel’s consequence obscuring the work OpenAI and the humans moving location person for the AI systems they designed, deployed, and grounded to contain. “Follow the incentives,” he said. Psychologist and influential AI skeptic Gary Marcus made a akin statement successful a Substack blog of his own, claiming anthropomorphic connection “distracts from the existent problems astatine hand.” And it’s each successful OpenAI’s liking to support that communicative going, he argues: “The ungraded is the inept in-house information astatine OpenAI. And the marketing. With gullible podcasters amplifying the PR.”

In X posts responding to his galore critics, Patel has defended his prime of words. Part of it is practical: location is nary evidently neutral vocabulary to picture what these agents did. Either we usage acquainted connection of intentions, goals, and collaboration and consequence implying excessively much, aliases trim everything to codification and usage cold, mechanical connection that risks stripping distant important elements of what we see. “Many group look to judge that if alternatively of a ‘civilization’, I had called them a ‘swarm of matrices’, location wouldn’t beryllium a problem worthy worrying about,” Patel said.

Complicating matters further is that the anthropomorphic connection doesn’t only travel from Patel, aliases moreover from the humans studying the agents. Terms for illustration “sacrifice,” “honor,” and “coalition” characteristic successful the agents’ transcripts. Google AI interrogator Neel Nanda argued that “anthropomorphic connection is reasonable” successful specified circumstances.

Doublespeak it is, then. Human-laced connection risks saying excessively overmuch astir what these systems are, and coldly mechanical connection risks saying excessively small astir what they tin do. Until we find connection tin of capturing both, the 2 contradictory ideas whitethorn simply person to coexist.

Follow topics and authors from this communicative to spot much for illustration this successful your personalized homepage provender and to person email updates.

More