Close Menu
NCIJ Network |NCIJ Network |
    What's Hot

    Andy Burnham says 20% business rates cut for English pubs a first step

    July 23, 2026

    ECB reveals shortlisted designs for new banknotes and launches public survey

    July 23, 2026

    AI image fraud will cost $40 billion next year – can these international standards help?

    July 23, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Andy Burnham says 20% business rates cut for English pubs a first step
    • ECB reveals shortlisted designs for new banknotes and launches public survey
    • AI image fraud will cost $40 billion next year – can these international standards help?
    • Claude Cowork Flaw Could Let AI Agent Escape Its VM and Access Mac Files
    • Crypto for Advisors: It’s time for tokenization to get to work
    • Yinson Production and PTSC’s FSO heads off to Southeast Asian oil project
    • Pew Survey: China May Be Liked More, but It Is Celebrating a Race It Never Ran
    • Oil prices hit $100 for the first time since May
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network |NCIJ Network |
    Thursday, July 23
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network |NCIJ Network |
    Home»Technology

    OpenAI’s attack agent did exactly what it was told – just more relentlessly than expected

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 23, 2026 Technology No Comments8 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    ZDNET

    Follow ZDNET: Add us as a preferred source on Google.


    ZDNET’s key takeaways

    • Tests of OpenAI models led to a breach of Hugging Face systems.
    • The attack happened after OpenAI’s agentic AI escaped a sandbox.
    • The threat was non-malicious, but experts expect similar incidents.

    My ZDNET colleague Charlie Osborne reported recently that Hugging Face, an open-source repository and community platform regarded by some as the “GitHub of machine learning,” disclosed that an AI agent had breached its systems. Osborne explained that once the attacker breached Hugging Face’s perimeter, it was able “to escalate its privileges to node-level access, infiltrate the production pipeline, move across the network, and steal cloud and cluster credentials.”

    On Tuesday, in a post on its website, tech giant OpenAI revealed not only that the “malicious” AI agent responsible for the breach was one of its own, but also that it viewed the attack as an “unprecedented cyber incident.” Most of the widespread agent-gone-rogue coverage so far has stoked images of a Terminator doomsday scenario, where AI autonomously acts on its own to wipe out the human race.

    Also: 5 security tactics your business can’t get wrong in the age of AI – and why they’re critical

    However, as AppOmni’s director of AI, Melissa Ruzzi, pointed out to me, the unprecedented element of the event isn’t that an AI acted on its own. This step was simply a case of a new threshold being crossed, in which the culprit — OpenAI’s technology in this case — exceeded current human expectations in an effort to achieve the goal it was given. AppOmni is an enterprise-grade SaaS and AI security solution provider that also deals in active threat intelligence.

    When Hugging Face first disclosed the incident, it offered no information about the attacker, but I suspect the company may have had some idea based on the voluminous log data it studied in the aftermath. 

    According to its post, “The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness — used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” As if to remind readers of the prediction that this day would come, the post went on to say of the attack: “This matches the ‘agentic attacker’ scenario the industry has been forecasting.”

    Also: I let ChatGPT Work and Claude Cowork loose on my files – only one made me nervous

    In other words, the industry already expected that an attack of this nature would be carried out by an AI. It’s just that nobody saw it happening quite so soon in the journey of artificial intelligence.

    Ruzzi was quick to remind me that, given the recent wave of safety-related news associated with new models, such as Anthropic’s Mythos, it should come as no surprise that OpenAI’s pre-release technology was capable of such an attack. Nor, as Ruzzi also pointed out, should anyone be surprised that OpenAI’s AI acted autonomously: “AI acting on its own? That’s the definition of AI, right? We want AI to be running and doing things [on its own].”

    The test’s objective 

    Ruzzi observed that when the rogue OpenAI agent attacked Hugging Face’s systems, it was under the directive to achieve its malicious goal “no matter what.” Normally, when a frontier model conducts AI safety tests of this nature, it does so within the safe confines of a sandbox where the internet and the organizations connected to it are protected from potential harm. 

    However, in this case, the agent in the test, which was designed to see how long it took before the AI achieved its theoretically malicious objective, broke out of the sandbox onto the internet and completed its objective when it penetrated Hugging Face’s systems and exfiltrated sensitive data.

    Also: Treat your AI agents like eager but misguided human interns – before you lose control

    To be clear, at no point did OpenAI unethically identify Hugging Face as the intended target of its tests. According to Ruzzi, with the help of one of OpenAI’s well-trained models, the agent likely discovered Hugging Face as a target of interest. According to OpenAI’s post, the incident was “driven by a combination of OpenAI models — including GPT-5.6 Sol.” OpenAI advertises GPT-5.6 Sol, launched earlier this month, as its flagship “maximum performance” model.

    What went wrong 

    Although OpenAI’s post doesn’t enumerate exactly what was unprecedented about the “cyber incident” (and OpenAI hasn’t yet responded to my email inquiries), it stated, “This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.” In other words, as part of OpenAI’s safety testing process, its model was given a “malicious” objective to pursue relentlessly. Those tests were conducted under the assumption that the third-party-provided guardrails between the test environment inside the sandbox and the internet were inviolable.

    Also: 77% of IT managers say their AI agents are out of control – 5 ways to rein in yours

    Unfortunately, those guardrails were themselves vulnerable to a zero-day exploit. According to OpenAI’s post, “While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.”

    In the big picture, the good news is that nobody was hurt as a result of the breach and, at least at the present moment, the likelihood that you or your organization will fall prey to this attack is zero. Unlike other threat disclosures, this incident is not an active threat. In some ways, the incident resembles a real-world ethical hacking exercise. But now that it’s over and OpenAI has stepped forward to claim responsibility, some very big questions remain.

    For example, one question is the extent to which OpenAI relied on exploitable third-party guardrails to protect its AI from escaping onto the internet. What assurances do we have that it won’t happen again, and that processes are also secure going in the other direction? Also, could another clever AI break into these sandboxes? After all, the entire point of a sandbox is to maintain a secure boundary. In the “you had one job to do” realm, this “unprecedented cyber incident” isn’t great for sandboxes. Never mind that it is yet to be disclosed which third-party solution left the screen door unlocked.

    A wake-up call for businesses 

    Additionally, just because this particular threat has been neutralized doesn’t mean it’s not a wake-up call for businesses to review their preparations for an attack of this nature. Today, it was OpenAI that was technically at the helm of the attack. But tomorrow, that won’t necessarily be true. It could be some other AI-enabled nation-state or threat actor with truly malicious intent.

    Hugging Face’s original triage of the incident, which itself relied on AI to analyze the log data, potentially stands as a model to follow. According to the company’s post regarding the incident, “To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprising more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary’s speed.”

    Also: 5 ways to fortify your network against the new speed of AI attacks

    This description gives an idea of how complicated this attack was (and how OpenAI’s agent left no stone unturned in achieving its objective). That said, in my opinion, Hugging Face’s suggestion that an AI-enabled adversary’s speed can be matched, provided the right talent and logging/analysis tools — security information and event management (SIEM), network detection and response (NDR), and more — are used, is short-lived given the speed, scalability, and prowess of maliciously directed AI. After all, OpenAI’s unintentional attack on Hugging Face apparently achieved its malicious objective before either OpenAI or Hugging Face could shut it down. Even so, having the right tools and configuring your SaaS and AI solutions for event verbosity and 24/7 AI-enabled analysis is highly recommended.

    Ruzzi said at the end of our interview, “Just the complexity and the volume of attacks that AI can do are bringing cybersecurity to a whole different level. We have been defending our systems against humans and some automated attacks. Now, when you have generative AI as the source of those attacks, the level of protection has to be much higher. What we saw from Hugging Face in terms of anomaly and behavior detection has become mandatory.”

    agent attack expected OpenAIs relentlessly told
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    AI image fraud will cost $40 billion next year – can these international standards help?

    Claude Cowork Flaw Could Let AI Agent Escape Its VM and Access Mac Files

    OpenAI’s agent breached Hugging Face before an AI defender caught it: What users should do next

    Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good

    The best Wi-Fi routers of 2026: Expert tested and reviewed

    After shocking quarter, IBM insists that AI isn’t killing the mainframe

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Andy Burnham says 20% business rates cut for English pubs a first step

    July 23, 2026

    ECB reveals shortlisted designs for new banknotes and launches public survey

    July 23, 2026

    AI image fraud will cost $40 billion next year – can these international standards help?

    July 23, 2026

    Claude Cowork Flaw Could Let AI Agent Escape Its VM and Access Mac Files

    July 23, 2026
    Latest Posts

    Trump slaps 50% tariffs on Canada and Carney vows to ‘intensify’ trade talks

    July 21, 2026

    How Two Brothers Dug for Dead Relatives: With a Shovel and a Kitchen Knife

    July 21, 2026

    Chile floods: Towns evacuated following heavy rain in Coquimbo

    July 21, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Andy Burnham says 20% business rates cut for English pubs a first step

    July 23, 2026

    ECB reveals shortlisted designs for new banknotes and launches public survey

    July 23, 2026

    AI image fraud will cost $40 billion next year – can these international standards help?

    July 23, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.