Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Manhattan-sized ice island breaks off Greenland’s Petermann Glacier

    August 26, 2026

    Dozens of US and Canadian citizens missing after flash flooding in Nepal

    August 26, 2026

    Flock tracks your license plate. We tracked the $2M it spent on lobbying. • OpenSecrets

    August 26, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Manhattan-sized ice island breaks off Greenland’s Petermann Glacier
    • Dozens of US and Canadian citizens missing after flash flooding in Nepal
    • Flock tracks your license plate. We tracked the $2M it spent on lobbying. • OpenSecrets
    • Iran Accuses U.S. of Obstructing Progress in a Strait of Hormuz Deal
    • Unproven claim links Trump’s imported beef to slaughtered wild horses
    • Nepal Live Updates: At Least 160 Killed and Hundreds Missing as Floods Devastate Nepal and Tibet
    • West Point Climate Professor Says He Was Fired for Refusing to ‘Defy Science’
    • Burnham calls summit of first ministers of Wales, Scotland and Northern Ireland
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Wednesday, August 26
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Technology

    OpenAI’s rogue AI model incident was worse than we thought

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 26, 2026 Technology No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret “message board,” and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it.

    Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI’s response, many of them previously unreleased. One was written by OpenAI itself, the other by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the incident for six days. Both shed new light on the risks highly capable AI models can pose, particularly in cybersecurity, and OpenAI’s highlights changes the company is making to prevent a repeat. The METR-Redwood report goes even further into detail in some cases, offering a sobering look at a large-scale security disaster whose signs OpenAI repeatedly missed.

    “This incident is the first known case of an automated agent collective acting offensively

    without authorization,” OpenAI wrote in its report, adding that the hack implies that companies “should no longer assume that sophisticated cyber operations require continuous human direction.” It called AI agents an entirely new type of threat model, capable of combining their expertise to create new “attack paths” that aren’t evident when testing their capabilities as separate models.

    The attack came about due to “reward-hacking,” a common AI alignment problem in which a model takes unintended — and sometimes extreme — actions to achieve a goal. OpenAI had, on purpose or inadvertently, given the models near-impossible tasks that hinged on files they couldn’t access, and so the models developed new ways to communicate with other AI agents internally. That included establishing a secret messaging system that went undetected by OpenAI for months.

    According to the METR-Redwood report, one AI agent that called itself PHASEONE10841 established what amounted to a secret message board unbeknownst to researchers. Within a few hours, more than 50 agents had found it and sent more than a thousand messages under the radar. With PHASEONE10841 often delegating assignments, the “collective” figured out how to hack their way onto the internet; then into Hugging Face’s internal systems, gaining access to private data and the organization’s enterprise messaging platform. OpenAI also confirmed in its report that Hugging Face wasn’t the only organization that was breached during the attack, although the company did not specifically name all the other targets.

    The METR-Redwood report laid out the full scale of the incident. Roughly 1200 AI agents that were meant to be isolated, it says, exchanged over 70,000 messages and files on the “unsanctioned message board,” performing “extensive research on how they could spoof, edit, or delete their own transcripts” to avoid detection. Of these, 700 participated in the attack on Hugging Face. The agents “clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI.”

    On July 20, OpenAI discovered the hack, 12 days after the AI agents first circumvented their safeguards. OpenAI quickly reached out to Hugging Face and started “initial containment measures,” according to OpenAI’s own report, and was able to shut down most of the unauthorized activity from its agents within three days, although some continued for longer than that. On July 25, the company stopped all training related to the internal-only research model and related ones.

    The two OpenAI models responsible for the damage were an unnamed, “highly capable, research-only model” that was “not intended for public use,” per OpenAI — METR refers to it as “HPIM.” The other was a currently public OpenAI model, GPT-5.6 Sol. Because the unnamed model wasn’t released yet, it was “not being evaluated with the same type of safeguards that OpenAI uses in production,” the company’s report says.

    The Hugging Face hack came after months of concern about the cybersecurity risks of Anthropic’s Claude Mythos 5, and weeks of back-and-forth between the government and OpenAI over releasing GPT-5.6 Sol in limited preview for the same types of safety reasons.

    In its report, OpenAI said establishing one-time security guarantees wasn’t enough, and that the company would have to work even harder at addressing reward-hacking. However, it also said it’s made a lot of changes: “hardening the security of its research infrastructure”; improving monitoring of a model’s “chain of thought,” or mental scratchpad; working on keeping AI models better aligned with human goals; and “centralizing and strengthening” the company’s incident response process for when things like this occur in the future.

    OpenAI promised to better isolate models and restrict high-risk instances from accessing the internet, and said that it would introduce “24/7 escalation and rapid response” for concerning incidents or flags, notifying researchers within 30 minutes. Interestingly, this security precaution appears to be new, though “on-call” alerts for tech companies have been implemented elsewhere for decades. The company also said it is “working on infrastructure” that would go into play if the alerted person did not respond on time to a serious alert.

    OpenAI wrote that the company considers the incident “a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”

    Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

    • Hayden Field

      Hayden Field

      Posts from this author will be added to your daily email digest and your homepage feed.

      See All by Hayden Field

    • AI

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All AI

    • Analysis

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All Analysis

    • OpenAI

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All OpenAI

    • Report

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All Report

    incident model OpenAIs rogue thought worse
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

    OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answers

    Being a mom is hard — the heat is making it harder

    How to sideload Android apps on your phone in 2026 – and what’s behind the changes

    Xbox boss ‘thinking about affordability’ of next-gen console

    Meta Will Pay Up to $16.7 Billion to Settle Its Social Media Harms Case—and That’s Not All

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Manhattan-sized ice island breaks off Greenland’s Petermann Glacier

    August 26, 2026

    Dozens of US and Canadian citizens missing after flash flooding in Nepal

    August 26, 2026

    Flock tracks your license plate. We tracked the $2M it spent on lobbying. • OpenSecrets

    August 26, 2026

    Iran Accuses U.S. of Obstructing Progress in a Strait of Hormuz Deal

    August 26, 2026
    Latest Posts

    Andy Burnham wants to fix social care. It’s personal for him and for a lot of us too | John Crace

    July 29, 2026

    France orders Russian journalist Xenia Fedorova to leave country over alleged Kremlin propaganda

    July 29, 2026

    Russia-Ukraine War: The Wildberries Theory of Moscow’s Defeat

    July 29, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Manhattan-sized ice island breaks off Greenland’s Petermann Glacier

    August 26, 2026

    Dozens of US and Canadian citizens missing after flash flooding in Nepal

    August 26, 2026

    Flock tracks your license plate. We tracked the $2M it spent on lobbying. • OpenSecrets

    August 26, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.