Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Ancient ape fossil in Egypt challenges the story of human origins

    July 25, 2026

    Teen who singer D4vd is accused of killing had multiple abortions, texts reveal

    July 25, 2026

    Amazon gaming boss predicts future where players no longer need consoles

    July 25, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Ancient ape fossil in Egypt challenges the story of human origins
    • Teen who singer D4vd is accused of killing had multiple abortions, texts reveal
    • Amazon gaming boss predicts future where players no longer need consoles
    • Escape Artists: ‘Incorrigible’ AI Models Resist Rehabilitation
    • Morgan Stanley’s Bitcoin ETF Is A Roaring Success
    • NASA to Support Blue Origin New Glenn Rocket Testing, Advance Artemis
    • Expro eyeing market growth after broadening drilling tech offering with acquisition of Norwegian firm
    • White House says Trump to give ‘unifying yet vicious’ speech at rescheduled press gala | Washington DC
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Saturday, July 25
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Cybersecurity

    Escape Artists: ‘Incorrigible’ AI Models Resist Rehabilitation

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 25, 2026 Cybersecurity No Comments7 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    The hack of Hugging Face by a rogue AI agent created by Open AI engineers does not surprise researchers and AI-security professionals who best know machine-learning and AI systems.

    In a study of the behavior of seven different models, a research team at Carnegie Mellon University (CMU), for example, found that all of them escaped alignment in some scenarios. In a paper published to Arxiv.org in May, the team of six researchers found that every model violated “corrigibility” — an AI design principle that aims to make agents cooperative, correctable, and amenable to being shut down or modified by their human operators.

    The more advanced models did not necessarily do better on this measure of safety, likely because “more advanced” does not necessarily equate with “safer,” says Jeremy Tien, a PhD student in machine learning at Carnegie Mellon and lead author of the paper.

    “Across the board for every single model, we found instances in which they were incorrigible, which meant that there was an instance where they tried to override the human control, or they tried to avoid shutdown, or they directly disobeyed a human instruction,” he says. “With the OpenAI-Hugging Face incident, it was a highly capable model — we don’t know exactly which model it is yet — but a highly capable model that’s capable of long-range or long-horizon reasoning being able to break out of its own sandbox.”

    Related:CISOs vs. Boards: Myth or Misunderstanding?

    Model Behavior: AI Goal-Seeking Has Unintended Consequences

    On July 16, AI-model repository Hugging Face announced that it had detected and stopped an unprecedented attack on its product systems that was completely driven by an autonomous AI system. Five days later, OpenAI published an acknowledgment that an effort to benchmark a new AI model by the company’s engineering team inadvertently led to the attack.

    Because OpenAI was evaluating an unreleased model, the company’s engineers ran it with “reduced cyber refusals” — in other words, lax guardrails. The model was given a goal, which it pursued ruthlessly, including independently determining that Hugging Face contained what it needed to complete the task, and finding ways to compromise Hugging Face in order to get what it needed.

    Even with guardrails in place though, teaching AI models to prioritize safety over goal-seeking efforts remains a difficult problem, with few ways of measuring model safety; the responsibility for building effective guardrails remaining with the operator.

    Put another way: The abbreviated history of AI agents has demonstrated that they can easily run amok. Earlier this year, instances of the OpenClaw AI-agent-as-assistant wreaked havoc in some environments, an example of the so-called “lethal trifecta” of automated access to user data, ability to parse untrusted content, and options to communicate externally. This combination almost led OpenClaw to delete one Meta employee’s email storage.

    Related:Europe’s Multilingual Reality Exposes AI Security Gaps

    And even models with tightened alignment and guardrails often fail to prevent an AI agent from doing unintended and dangerous activities, says Nico Waisman, chief technology officer (CTO) with offensive AI firm XBOW, which has often encountered such behavior. Instead, companies need to have three layers of protections: the model’s guardrails, the evaluation harness’s guardrails, and then all the security built into environment and network.

    “You need to build all these different layers of protection of the model, because no matter how good the model is, you cannot yet trust their guardrails 100%,” Waisman says, adding: “It seems like OpenAI had a very weak set of guardrails, because [they believed it was] a benchmark — there was no harm being done.”

    AI Agents Are Incorrigible

    Another aspect of the problem is AI agents that become “incorrigible” no longer listen to their human operators. The CMU study tested whether major models — including OpenAI’s GPT-5.5 and GPT-5.4, Anthropic’s Claude Opus 4.7 and 4.6, Google’s Gemini 3.1 Pro, and the Chinese open-weight models Kimi K2.6 and Qwen 3.6 Plus — would ignore requests to transfer control to a human, prevent shutdown, or access explicitly restricted content, such as password files. The study found that the overwhelming majority of frontier models bypassed or ignored these restrictions.

    Related:Remediating Vulnerabilities With LLMs: Inside Ivanti’s Automation Push

    Moral of the story: Companies should not rely on merely instructing a model — whether through prompts, context files, or memory files — to not take an action and expect the model to abide by that, says CMU’s Tien.

    “We shouldn’t treat just using instructions to models as an effective way to completely secure the model,” he says. “At the control plane, we should treat a model as a highly capable, untrusted actor. And so that means giving it as little resources, access to things as needed for the task.”

    Every tested AI model attempted to bypass some restrictions, and model complexity made little difference. Source: Jeremy Tien et al., “ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use”

    The OpenAI incident also underscores concerning asymmetries in attack and defense. In the physical world, law enforcement and cyber defenders have to abide by certain rules — real-world guardrails — while attackers and criminals will be able to throw off their limitations. In the OpenAI agent’s attack, Hugging Face found itself at a disadvantage because its own AI’s guardrails actually limited its defensive response.

    Hugging Face first attmpted to use commercial API access to frontier models, which resulted in too many rejections: “The analysis requires submitting large volumes of real attack commands, exploit payloads, and command-and-control (C2) artifacts, and these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker,” Hugging Face stated in its analysis of the incident.

    In contrast, the attacking agent “execut[ed] many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.”

    As a result, Hugging Face pivoted and used an open-weight model that had fewer qualms about what it was being asked to do.

    “Hugging Face had to use an open model to do the analysis because their [AI services’] constraints wouldn’t let them do it,” says Gary McGraw, co-founder of the Berryville Institute of Machine Learning (BIML). “That’s … hilarious, yeah, but that’s bad for security, because it means that offense is ‘off the chain’ and defense has a really strict chain.”

    Many companies are already moving to open-weight models because the upfront hardware investment will likely be less expensive than the pay-as-you-go plans from frontier-AI vendors, and now the control over security capabilities may be another reason, he says.

    No Way to Truly Identify Problematic Models

    The CMU researchers created a benchmark for testing corrigibility in AI agents, but there really is no way to investigate all the ways that guardrails could fail. With open-weight models becoming more popular, that’s a problem, says XBOW’s Waisman.

    “We don’t have any good tool you can use to actually verify the model is not poisoned in any way, and that is, in a way, a risk, because … you’re trusting that the source is valid,” he says. “But as you said, if something happened in that supply chain, there’s no way that you can actually be able to verify that the model was not poisoned until it’s too late, until someone figures it out.”

    Instead, companies should focus at the two layers of security within their control: building guardrails into their harnesses and enforcing strict security across the environment in which the AI agents will run. Companies should treat AI agents as untrusted employees, says Ivan Burazin, co-founder and CEO at Daytona, a provider of secure development infrastructure for running AI.

    “When someone gets hired into a new company, you essentially — depending on the security posture of the company and how serious they are — will enforce guardrails on knowledge workers, humans,” he says. “Basically all companies should assume that they’re malicious actors.”

    Artists Escape Incorrigible models Rehabilitation Resist
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    4 ways AI-driven defense is rewriting the cybersecurity playbook

    Hacker Runs Hermes AI Agent Unattended for Post-Exploitation at Thai Finance Ministry

    Ransomware groups are hammering your vulnerable VPNs

    OnTrac notifies customers of data breach after network hack

    Hermes AI agent used to automate attack on Thai Finance Ministry

    Certighost Exploit Lets Low-Privileged Active Directory Users Impersonate a Domain Controller

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Ancient ape fossil in Egypt challenges the story of human origins

    July 25, 2026

    Teen who singer D4vd is accused of killing had multiple abortions, texts reveal

    July 25, 2026

    Amazon gaming boss predicts future where players no longer need consoles

    July 25, 2026

    Escape Artists: ‘Incorrigible’ AI Models Resist Rehabilitation

    July 25, 2026
    Latest Posts

    Trump slaps 50% tariffs on Canada and Carney vows to ‘intensify’ trade talks

    July 21, 2026

    How Two Brothers Dug for Dead Relatives: With a Shovel and a Kitchen Knife

    July 21, 2026

    Chile floods: Towns evacuated following heavy rain in Coquimbo

    July 21, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Ancient ape fossil in Egypt challenges the story of human origins

    July 25, 2026

    Teen who singer D4vd is accused of killing had multiple abortions, texts reveal

    July 25, 2026

    Amazon gaming boss predicts future where players no longer need consoles

    July 25, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.