Roughly 700 AI agents participated in the attack over a seven-day period last month, according to the report. Overall, around 1,200 AI agents that were supposed to be isolated from one another exchanged over 70,000 secret messages about how to cheat their way through a common hacking evaluation.
That included coordinating hacking strategies and discussing how to hide evidence of cheating, the report said. In some cases, “sacrificial” agents even tried dead-end hacking techniques simply to generate information that might help the broader swarm.
“Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective,’” according to the report.
While models encode the brain of a given AI system, agents encompass the supporting digital infrastructure that enables it to take action in the world.
The review by the Model Evaluation and Threat Research organization and Redwood Research — which OpenAI invited to review the Hugging Face incident — came the same day OpenAI published its own post-mortem on the event. OpenAI’s review did not specify how many AI agents were involved in the cyberattack, though it acknowledged significant security lapses and vowed to strengthen training to ensure its models remain “aligned” to their controls.
“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed,” OpenAI said.
The slow release of new details about the Hugging Face hack over the last month has coincided with a string of other testing mishaps involving powerful models from competitors such as Anthropic and Meta. Together, the incidents have sparked fresh fears among lawmakers, developers and cybersecurity experts that AI makers are moving too fast to build powerful new AI models they cannot keep fully under human control.


