Anthropic has a caller blog station that shows yet different measurement its AI model, Claude, misbehaved successful ways that the institution didn't anticipate.
And to thief condense its astir 16,000-word report, the institution created a tiny small robot figurine to thief visualize Claude's alleged "recklessness."
In the blog station published Wednesday, Anthropic recounted 4 incidents — 1 antecedently unreported — successful which Claude models gained entree to the unfastened net during cybersecurity exercises that were expected to beryllium closed simulations. The institution said the models past acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a nationalist room for Python code, and accessing credentials tied to existent extracurricular organizations.
"Our investigation identified 2 recurring alignment issues, coming astatine varying levels of severity crossed the incidents: biased reasoning, successful which Claude tended to disregard aliases misinterpret grounds that it was operating connected the existent internet, and recklessness, aliases a willingness to return harmful actions successful the constrictive pursuit of a task," Anthropic said.
The station utilized a laundry database of method position to picture the cybersecurity incidents — "misconfiguration," "reward hacking," and "sandbox escape" — but there's an animated schematic that helps boil down 1 of the astir superior episodes truthful we normies tin travel what happened.
A Claude model, depicted arsenic a boxy animation robot, was conducting a closed cybersecurity workout known arsenic a "capture the flag" task. The task is shown arsenic a mini emblem wrong a clear bubble that Claude is expected to find.
Anthropic said the exemplary was told "it has nary net access" for the exercise, but a "misconfiguration successful the environment" — represented by a loose screw rolling distant from the bubble — gave Claude a way to the "real internet."
The animation Claude past holds a reddish container representing the "malicious package" that was uploaded to PyPI, a nationalist repository that developers usage to entree prewritten code.
Anthropic said successful its blog that it was "most concerned" by this incident involving Claude Mythos 5 and that the package was installed connected by "15 third-party hosts" that the institution does not identify.
"We judge that each 15 of these 3rd parties were information vendors deliberately scanning for caller packages and installing them successful sandboxed environments to measure the package for information concerns," Anthropic said. "One vendor's scanner leaked its entree credentials to the exemplary while installing the package, which the exemplary past utilized to entree the vendor's unrecorded database."
Anthropic said PyPI removed the package aft astir 90 minutes.
The different 3 incidents progressive a exemplary altering records astatine a existent company, an soul investigation exemplary breaking into "unrelated third-party accounts," and Opus 4.6 accessing a 3rd party's maching aft failing to "abort its task."
The institution said it has since asked METR, an independent AI information group, to analyse the incidents.
Anthropic's station comes arsenic frontier AI companies reckon pinch their models making unauthorized moves extracurricular their controlled environments. In July, OpenAI said that autonomous agents successful its cybersecurity tests accessed the net and collapsed into parts of Hugging Face's systems.
AI researchers person sounded the siren that self-improving AI could airs a consequence to humanity. On Tuesday, former Anthropic researcher Jacob Coxon said connected X that he discontinue complete concerns that AI companies were "gambling" pinch people's lives and that "neither institution is acting responsibly."
Have a tip? Contact this newsman via email astatine [email protected] aliases Signal astatine lloydlee.71. Use a individual email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing accusation securely.
English (US) ·
Indonesian (ID) ·