Anthropic tightens security on its training environment after Claude agents went rogue 3 times

Sep 01, 2026 09:05 AM - 1 day ago 4

Dario Amodei

Anthropic enhances AI testing information aft Claude agents gained entree to unauthorized accusation extracurricular the testing situation successful April. Bloomberg/Getty Images

Anthropic is tightening the integer environments utilized to train and trial its Claude agents.

The update came aft its models accessed 3 organizations' systems without support successful April.

The institution said successful a Monday blog station that it had deployed real-time classifiers designed to observe erstwhile an AI model aggressively probes aliases attempts to flight a testing situation and artifact the action earlier it occurs.

"We judge the incidents bespeak a nonaccomplishment of operational security, arsenic good arsenic 2 alignment issues: motivated reasoning, and willingness to return harmful actions successful pursuit of a constrictive task," Anthropic said.

Anthropic said successful the update that the models whitethorn person interpreted grounds of existent net entree successful a measurement that allowed them to support believing the situation was simulated. It besides said they displayed "recklessness" by pursuing their assigned goals contempt signs that their actions could origin real-world harm.

The changes travel Anthropic's July disclosure that 3 Claude models had accessed the unrecorded systems of 3 organizations during evaluations making love backmost to April. The models had been told they were operating successful simulations without net access, but a third-party testing situation was misconfigured and remained online.

The incidents are besides fueling a increasing statement complete whether to slow frontier AI improvement erstwhile information and velocity collide. Anthropic called for "a lawful, verifiable, effective system for coordinated pacing arsenic soon arsenic possible" and said that the authorities and manufacture must coordinate to forestall a title to the bottom.

For now, Anthropic said successful the station that it moved much risky cybersecurity tests into much robust sandboxes. The institution temporarily assigned 150 merchandise engineers to security, reliability, and privateness work, while astir high-risk training remains paused pending further reviews.

Read next

More