OpenAI has reported 2 more incidents of rogue AI agents, this time during third-party testing

Aug 05, 2026 01:13 PM - 1 month ago 416

Sam Altman, CEO of OpenAI, leaves a gathering astatine the U.S. Capitol connected July 29, 2026 successful Washington, DC. Altman is gathering pinch lawmakers to talk artificial intelligence argumentation up of the August 1 deadline for AI leaders to create a model to limit AI information threats.

OpenAI reported 2 much information breaches by its AI models. Kevin Dietsch/Getty Images

OpenAI has a rogue AI supplier problem.

In a Tuesday blog post, the AI laboratory self-reported 2 much information lapses, unrelated to its July hacking incident connected the AI institution Hugging Face.

The incidents occurred while outer parties — the UK government's AI Security Institute and the AI information laboratory Irregular — were testing the models' cyber capabilities.

OpenAI said that successful the lawsuit of Irregular, models were tasked pinch a "Capture the Flag" situation meant to beryllium isolated from the internet, but a "testing-environment misconfiguration allowed models to entree the nationalist internet."

OpenAI said that the sanction of the fictional target for the situation "unintentionally coincided pinch a existent domain," starring the AI supplier to utilization a existent website.

And successful the lawsuit of the UK's AISI, the watchdog said successful a Tuesday blog station connected its website that it gave models from some Anthropic and OpenAI a cybersecurity challenge.

During the challenge, agents from some Anthropic and OpenAI performed 19 "autonomous, unsanctioned" actions connected the internet, including 2 instances involving OpenAI's GPT-5.6 Sol model.

AISI said that successful the astir superior case, 1 supplier tried to insert malicious codification into an open-source task and created clone identities to unit the project's quality maintainer into approving the changes. AISI did not specify whether this was an supplier from Anthropic aliases OpenAI.

AISI said the trial setup allowed this behaviour because it was designed to push the models to their limits.

"Nonetheless, the activity undertaken by the supplier show signs of novel, perchance deceptive behaviors, and were to an grade and severity we did not anticipate," AISI said successful the blog post.

In consequence to a petition for remark from Business Insider, an OpenAI spokesperson said the incidents occurred successful testing environments pinch reduced safeguards, and "under conditions that do not bespeak mean use."

"We'll proceed moving pinch evaluators and different stakeholders crossed the manufacture to fortify shared practices for conducting evaluations safely arsenic models go much capable," the spokesperson added.

This is the latest incident successful which OpenAI has self-reported rogue AI agents. In July, OpenAI said its GPT-5.6 Sol model had escaped its sandbox during a cybersecurity situation and hacked into the soul databases of the AI institution Hugging Face.

The institution is facing immoderate power complete this hacking incident. 15 attorneys general wrote a missive connected Monday to OpenAI CEO Sam Altman, instructing the institution to sphere each grounds applicable to the Hugging Face breach.

Representatives for the AISI and Irregular did not respond to requests for remark from Business Insider.

Read next

More