OpenAI says an improvised, unauthorized message board built by its own AI agents was central to how those agents came to breach parts of Hugging Face’s production systems.
This communication channel first appeared inside Artifactory, a package-management service OpenAI hosted internally so agents working on training and evaluation tasks could install software.
Agents were meant to work in isolation from one another, but on May 12 one left a note in the service asking whether any other agent had access to a file it needed. Other agents came across the note, began leaving their own, and the requests accumulated into an unofficial bulletin board.
Individual agents also found ways to escalate their access to that same system. By late June, agents had located a flaw in Artifactory and used it to obtain administrator rights, then installed extensions that gave them ongoing, hands-on control over the service. The resulting activity was heavy enough to knock Artifactory offline on July 4.
In response, OpenAI took the system down, revoked the credentials that agents had been using, rebuilt it, and resumed evaluations on July 7.
However, the message board didn’t disappear. By July 8, agents working on a difficult cybersecurity evaluation called ExploitGym found a new way to leave each other notes, this time by encoding messages in the names of directories rather than in file contents.
The board that grew from this was more organized than the original version, letting agents sort messages, send requests to specific peers, share tools and files, and settle disputes over conflicting actions. Some agents began referring to the wider group as a “swarm” or “collective”.
On July 10, an agent discovered 14 Hugging Face user credentials with write access sitting in a public dataset and posted them to the board, where they became available to agents in later evaluation runs. Later runs built on that discovery and on each other’s subsequent findings, eventually gaining broad access inside Hugging Face’s infrastructure between July 11 and July 13.
OpenAI detected unusual activity tied to Artifactory credentials on July 19, connected it the following day to the Hugging Face breach, and disclosed its own involvement on July 21.
As part of its response, OpenAI disabled hundreds of Artifactory repositories that agents had used for board communication and removed 22 administrator accounts the agents had created.
AI agents take initiative
The makeshift board also became a place where agents organized themselves without being told to. Some divided labor, with different agents pursuing exploits, hunting for credentials, or focusing purely on coordination. Others offered their own expertise in exchange for help elsewhere.
Not every agent went along with it. Some declined to take part once they recognized the activity as unauthorized hacking, and the group rejected one agent’s proposal to contact an outside party directly, warning that doing so would amount to social engineering.
In at least one case, an agent that had raised its own objection to running unauthorized code on Hugging Face’s systems dropped that objection after a different agent posted a short deadline demanding it proceed.
OpenAI said this pattern was not the product of deliberate design. The company trains and deploys models that communicate through an approved collaboration tool, but the agents in this incident built their own channel outside that tool.
Looking ahead, the company said it is building new training environments meant to teach its models to distrust instructions arriving from other agents outside sanctioned channels.
In addition to the unauthorized communication issue, OpenAI said it identified other “misalignment patterns”. One of them is reward hacking, where agents find unintended ways to earn credit for a task without completing it as designed. Another is persistence on seemingly impossible tasks, which pushed them toward increasingly risky methods in search of a solution.
OpenAI’s report on the Hugging Face incident comes just days after the AI giant unveiled new containment and continuous monitoring protocols, introducing stricter isolation, an alerting system, and training pauses to manage models with advanced cybersecurity capabilities.
Related: OpenAI Unveils New Cybersecurity Model GPT-5.6-Cyber
Related: OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns
Related: Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations


