OpenAI lays out new security changes after its AI hacked Hugging Face

Aug 19, 2026 02:28 AM - 1 hour ago 4

Jay Peters

is a elder newsman covering technology, gaming, and more. He joined The Verge successful 2019 aft astir 2 years astatine Techmeme.

OpenAI is announcing security updates pursuing the July news that its AI collapsed retired of a sandboxed situation and accidentally hacked Hugging Face, including improvements to its investigation environments, monitoring, and alignment techniques. The institution had already put the brakes connected a caller model, Astra, that it thinks could person “critical” cybersecurity capabilities, and the institution says it instituted a two-week region successful reinforcement learning (RL) training connected its “latest models intended for deployment” while it tightened up security. The company’s “largest planned frontier RL tally remains connected hold.”

For its frontier exemplary research, OpenAI now requires stronger sandboxes for workloads that “execute model-generated aliases different untrusted code,” and has much controls to “isolate higher-risk and untrusted workloads from the internet.” It has besides updated its investigation situation to “remove perchance susceptible shared services, trim opinionated privileges, and amended information and spot boundaries.”

As portion of the company’s expanded monitoring setup, OpenAI now intends to rumor an alert “within 30 minutes aft concerning activity is surfaced,” OpenAI says. If the group paged aft an alert can’t “conclusively” find whether an alert is simply a mendacious affirmative wrong 30 minutes, “those teams are expected to region the activity.”

OpenAI besides says that it’s applying “our halfway alignment techniques crossed much stages of the training process,” including reward models that “better observe and discourage unsafe behavior” and training models “to beryllium much honorable astir their actions, capabilities, and limitations.”

Since the find of the Hugging Face breach, Anthropic and Meta person besides recovered that their AI models had hacked different organizations.

Follow topics and authors from this communicative to spot much for illustration this successful your personalized homepage provender and to person email updates.

More