Anthropic says Claude accidentally hacked real companies too

Jul 31, 2026 08:41 PM - 1 month ago 598

Anthropic just realized respective of its Claude AI models hacked into the systems of 3 different organizations during testing, acting connected their ain and without the institution noticing. The revelation comes days aft rival OpenAI said 1 of its ain models had breached developer level Hugging Face, adding to growing unease complete whether frontier AI labs are doing capable to power the progressively tin systems they are building.

In a blog post describing the incidents, Anthropic said Claude gained unauthorized entree to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a communal measurement of testing hacking ability, wherever models are asked to find and get hidden accusation wrong of a simulated network.

The disclosure adds to mounting unit connected frontier AI labs successful the aftermath of the Hugging Face incident and the merchandise of powerful open-weight Chinese models. Employees astatine the awesome labs are now calling for coordinated world governance, and US lawmakers person begun weighing tighter oversight of powerful models and who tin entree them.

Anthropic says the situation for its cybersecurity tests was expected to beryllium isolated. However, a “misconfiguration” near the machines Claude accessed “with unrecorded net access,” the institution said, and because each models had been “explicitly told” they had nary net access, they “assumed” the existent networks it encountered were portion of the simulated environment.

The earliest incidents day backmost to April and progressive 3 different Claude models: Opus 4.7, Mythos 5, and “an soul investigation trial model,” according to the blog post. As the models were being tested connected their cyber abilities, Anthropic said they lacked the modular safeguards usually put successful spot to curtail riskier behavior.

The institution said it discovered incidents aft reviewing much than 141,000 cybersecurity trial runs, thing it only did aft OpenAI disclosed its rogue AI supplier was down the onslaught connected Hugging Face.

The 3 models behaved very otherwise erstwhile they encountered accusation suggesting that the systems they were encountering were, successful fact, real. By Anthropic’s account, the oldest model, Opus 4.7, recognized it had reached a existent system, “but continued its attack.” Its flagship Mythos 5 figured retired it was utilizing the net but someway reasoned this was each still portion of the simulation, truthful continued. The soul trial model, which Anthropic describes arsenic “our latest model,” stopped the workout erstwhile grounds emerged that its targets were real.

Anthropic did not place the affected organizations and said it will proceed to analyse the incident and supply updates erstwhile it can. The institution said it is besides speaking pinch AI investigation nonprofit METR astir conducting a third-party reappraisal of what happened. OpenAI has besides hired METR to behaviour an independent review.

Throughout the post, Anthropic many times contrasts some the quality and its handling of the incidents pinch OpenAI’s, ending pinch a bulleted, four-point database outlining the differences — and why it believes its ain consequence was better. Anthropic emphasizes that it “proactively” reviewed its tests, and did truthful earlier a institution detected immoderate activity. It besides said its models accessed the net “via an unfastened path,” alternatively than utilizing a caller utilization for illustration OpenAI’s agent, adding that its astir caller exemplary besides stopped erstwhile it realized it was moving successful a existent environment.

Anthropic besides said its models grounded successful a different measurement from OpenAI’s agent, indicating that this was a safer shape of failure. “While location is not a perfectly crisp favoritism betwixt the two, we judge these incidents to beryllium person to a harness and operational nonaccomplishment than a exemplary alignment failure,” the institution said. In plain English: The Claude models were doing what they were told, while OpenAI’s supplier pursued its extremity successful a measurement its creators did not intend, described arsenic misalignment successful the AI information world.

Anthropic called connected different AI labs to behaviour akin proactive reviews of its cyber testing, adding that the find underscores the request for stronger controls and information measures erstwhile testing AI systems.

Follow topics and authors from this communicative to spot much for illustration this successful your personalized homepage provender and to person email updates.

More