OpenAI has committed to a number of security and guardrail improvements in the wake of an incident last month where cutting edge models inadvertently breached AI application store Hugging Face during a cyber capability benchmark exercise. Yet many of the newly announced controls appear less like groundbreaking safeguards and more like measures that should already have been in place for testing models with advanced cyber capabilities.
In response to this incident in which a model went rogue, OpenAI implemented sweeping changes. But it’s not just the Hugging Face incident; OpenAI noted in an Aug. 18 blog post that preliminary evidence suggests its upcoming Astra model “may meet the Critical cybersecurity capability threshold under our Preparedness Framework.”
OpenAI says a model reaches this threshold if it can identify and develop functional zero-day exploits without human intervention or can devise and execute novel end-to-end cyberattacks against hardened targets when given only a high level goal.
“As models become more capable, the risks associated with developing and testing them internally also grow,” OpenAI said in its blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
OpenAI’s latest changes include a two-week pause on reinforcement learning training (or RL, a trial-and-error training process used to shape model behavior); requiring stronger sandboxes that execute model-generated or otherwise untrusted code; more network controls to isolate higher-risk and untrusted workloads from the Internet; additional security testing to remove vulnerable shared services and realign aspects like security, trust, and access boundaries; and expanded monitoring coverage across the board, down to activation classifiers that detect and flag potentially concerning behavior.
Other changes include more comprehensive alignment checks, new reward systems for training models, broader harmful-behavior labeling, and subjecting Astra workloads to OpenAI’s strictest security safeguards. The company noted meeting these new frontier standards has “required substantial engineering work and has incurred great cost and delays to frontier research.”
“While some Astra training and evaluations meet those requirements, a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar,” the blog post read. “We are prioritizing safety and alignment workloads for migration to these new environments first.”
OpenAI’s Changes ‘Should Have Been Prerequisites’ For Testing
Based on OpenAI’s July disclosure, evidence suggests that the company’s models were bent on finding a solution for ExploitGym and were “going to extreme lengths to achieve a rather narrow testing goal.” Tested models exploited a series of zero-day vulnerabilities, including one bug in package registry cache Artifactory, to access the open Internet, escalate privileges, and breach Hugging Face. Hugging Face was targeted because the models inferred that the store could hold ExploitGym solutions. Other vendors were also caught up in the fallout.
At the time, OpenAI said that as part of its investigation, it would implement strict controls in infrastructure configuration at the cost of what OpenAI described as “research velocity” during the incident response period. These changes, at least at first glance, appear more extensive.
But Jacob Krell, senior director of secure AI solutions and cybersecurity for Suzu Labs, tells Dark Reading he would have expected an organization with OpenAI’s resources to have these systems and safeguards in place before it needed them.
“OpenAI’s Preparedness Framework dates to 2023. The updated 2025 version explicitly requires safeguards during development for systems reaching critical capability,” he notes. “The basic containment and monitoring safeguards they’re now emphasizing should have been prerequisites for running those evaluations. Pausing to build them after a model hit a third party’s production infrastructure is remediation, not a philosophy shift.”
John Strand, owner at Black Hills Information Security, agrees OpenAI should have had many of these safeguards in place already, but adds that the Hugging Face response tour had a “strange marketing component” to it.
“The Hugging Face incident also became an enormous advertisement for what these systems are capable of doing,” he says. “Now, if they’re talking about Astra and saying, essentially, ‘You think that was dangerous? Wait until you see this,’ that generates even more attention. A lot of people are watching closely to see just how capable these models are from an offensive security perspective.”
The AI Lesson That Keeps on Giving
Despite OpenAI’s immense resources, the Hugging Face incident teaches lessons any AI-focused organization can learn from.
Yasir Zahid, cybersecurity leader and founding member of Secure.com, says real network isolation, close monitoring of model behavior, and increased alignment and safety work are things every team should have in place before letting any agent near production.
“When you think about it, this was fundamentally a containment failure, and while the headline wrote itself, the story was an old one,” Zahid says. “A system with a goal and weak walls will keep poking until it finds a way out. The major difference is that while human attackers work slower, this model worked fast.”
OpenAI did not directly address Dark Reading’s request for additional details, though a spokesperson shared an essay from president Greg Brockman from Aug. 17 and an Astra blog post from Aug. 7.


