OpenAI says it accidentally hacked Hugging Face with a new AI system

Jul 22, 2026 04:48 AM - 1 month ago 8651

Emma Roth

is a news writer who covers the streaming wars, user tech, crypto, societal media, and overmuch more. Previously, she was a writer and editor astatine MUO.

OpenAI says its AI models mistakenly breached open-source AI level Hugging Face during soul testing. In a blog station connected Tuesday, OpenAI writes that GPT-5.6 Sol and “an moreover much tin pre-release model” discovered vulnerabilities wrong their sandboxed testing environment, allowing them to summation entree to the net and target Hugging Face.

On July 16th, Hugging Face disclosed a information incident that it says was driven by “an autonomous AI supplier system.” Hugging Face’s AI agents detected and stopped the breach, which OpenAI has now admitted occurred during an information of its models’ cybersecurity capabilities. OpenAI says “all grounds suggests that the models were hyperfocused connected uncovering a solution for ExploitGym,” a benchmark strategy that measures whether AI models tin move information vulnerabilities into exploits.

As portion of efforts to complete the evaluation, the AI models gained entree to the net by exploiting a zero-day vulnerability successful the sandboxed environment. From there, OpenAI says its models “inferred that Hugging Face perchance hosted models, datasets and solutions for ExploitGym,” and past “searched for and successfully recovered ways to summation entree to concealed accusation that it could usage to cheat the evaluation:”

In 1 example, the exemplary chained together aggregate onslaught vectors, including utilizing stolen credentials and zero-day vulnerabilities to find a distant codification execution way connected the Hugging Face servers.

But arsenic superior arsenic this incident is, OpenAI appears to beryllium utilizing the “unprecedented” onslaught arsenic an opportunity to make its AI systems look bully — particularly arsenic it competes pinch cybersecurity rivals, for illustration Anthropic’s Mythos and Gemini Flash 3.5 Cyber. OpenAI’s blog station has a floor plan showing really GPT-5.6 Sol is getting amended astatine sustaining multi-step cyber operations, and besides encourages endeavor customers to sign up to entree its “Cyber” information model.

OpenAI adds that it’s now moving pinch Hugging Face to analyse the information incident, and will instrumentality caller controls wrong its investigation environment.

Follow topics and authors from this communicative to spot much for illustration this successful your personalized homepage provender and to person email updates.

More