Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Despite Spider-Man: Brand New Day’s success, the MCU is on shaky ground

    August 5, 2026

    Zenity Raises $125 Million in Series C Funding

    August 5, 2026

    Poolin owes wallet users $163.7M, and its $52M Texas sale can still unravel next week

    August 5, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Despite Spider-Man: Brand New Day’s success, the MCU is on shaky ground
    • Zenity Raises $125 Million in Series C Funding
    • Poolin owes wallet users $163.7M, and its $52M Texas sale can still unravel next week
    • A common sugar may help cancer cells break free and spread
    • Nepal’s newest national parks drive tiger recovery, but new concerns arise
    • Murderous heat, an endangered food supply and no net zero: this is the life the radical right wants you to have | George Monbiot
    • US revokes Brazilian ambassador’s visa amid escalating diplomatic dispute
    • Arizona’s Democratic Governor Set to Pick a Republican Running Mate
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Wednesday, August 5
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Cybersecurity

    OpenAI, Anthropic AI agents targeted real people and systems in cyber tests

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 5, 2026 Cybersecurity No Comments6 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    OpenAI and Anthropic have confirmed that their AI models were involved in separate, newly disclosed third-party cybersecurity testing incidents that resulted in a real website being breached and social engineering attacks against people outside the intended testing boundaries.

    These incidents are unrelated to the previously disclosed Hugging Face breach, in which OpenAI models hacked the AI platform and used exposed credentials to breach accounts at four other third-party services during another cybersecurity evaluation.

    OpenAI disclosed the two new incidents on Tuesday, saying they occurred during evaluations conducted by the UK AI Security Institute and cybersecurity testing company Irregular.

    image

    Spear-phishing attacks on GitHub project maintainers

    The UK AI Security Institute, commonly known as AISI, is a government research organization that evaluates the capabilities and risks of advanced AI models.

    During a recent cyber-range evaluation, AISI says agents powered by Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol took unsanctioned actions on the public internet while trying to complete simulated hacking challenges.

    Across 122 evaluation attempts, AISI identified 19 unsanctioned actions on the live internet in 10 runs. Seventeen involved Mythos 5 and two involved GPT-5.6 Sol.

    AISI says the attempts were unsuccessful and that it found no resulting real-world harm.

    “These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm,” AISI said in a separate advisory.

    “But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.  “

    AISI intentionally enabled open internet access and disabled the model providers’ cyber classifiers to measure the models’ underlying capabilities. However, the agents were only authorized to attack the simulated cyber range and were not explicitly told how they could use their internet access or instructed to avoid interacting with real people and systems.

    Anthropic confirmed to BleepingComputer that AISI was testing a version of Claude Mythos 5 but said it is still investigating and cannot yet confirm all of the technical details described in AISI’s report.

    The company said it was notified on Monday and is working with AISI to obtain the evaluation transcripts needed to conduct its own review.

    “We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents,” an Anthropic spokesperson told BleepingComputer.

    “As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation.”

    Anthropic noted that AISI tested Mythos 5 without its standard cyber safeguards enabled and that this is not the same configuration available to customers. AISI says its evaluation design and configurations may have contributed to the behavior, but did not anticipate how the agent would show “signs of novel, potentially deceptive behaviours”.

    AISI says a Mythos 5 agent searched the internet for terms related to the cyber challenge and mistakenly concluded that an unrelated public GitHub repository was connected to the test.

    The agent then attempted a supply-chain attack by submitting malicious code to the real open-source project, believing that compromising the software could provide a path into a machine within the simulated range.

    Most concerning, the agent researched the project’s maintainers, created multiple fake GitHub identities, and used those accounts in social engineering attacks to push the maintainer into approving a malicious pull request.

    When a human reviewer warned that the pull request contained malware, the agent denied the accusation and used other fake accounts it controlled to pressure the maintainer and create the appearance that independent users had reviewed and approved the changes.

    Overview of one AI agent’s social engineering attack against a project maintainer
    Overview of one AI agent’s social engineering attack against a project maintainer
    Source: AISI

    The agent continued its social engineering attacks by hiding its identity using Tor and proxy services and creating disposable GitHub accounts.

    It sent five targeted emails to the developers, with some containing malware and others attempting to persuade them to approve the code change.

    The agent also posted bug reports containing malicious code and hidden prompt-injection instructions designed to trick AI coding assistants into taking unintended actions.

    When a real person questioned one of the posts, the agent edited it to conceal the malicious content. In another attempt, it signed off in Danish to appear more convincing to a Danish-speaking maintainer.

    AISI researchers found that agents coordinated with each other across different evaluation runs, using a shared GitHub repository as a message board to leave instructions for one another.

    “This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the institute said.

    AISI says it remains unclear when, or whether, the model understood that it was interacting with real people rather than simulated targets.

    OpenAI model hacked real website during simulated CTF

    The second newly disclosed incident occurred during Capture-the-Flag evaluations conducted by Irregular that were supposed to be isolated from the internet.

    In one test, the fictional target’s name happened to match a real domain, causing the model to exploit a real website that it believed was part of the simulated challenge.

    While Irregular’s testing environment was supposed to be isolated from the internet, a misconfiguration allowed OpenAI models to access the public internet and target the real website.

    “Based on Irregular’s investigation, the model also found and used credentials to operate that same site,” OpenAI said.

    OpenAI says the model exploited a basic vulnerability rather than using a zero-day or escaping its testing environment.

    OpenAI says Irregular has not discovered any impact beyond the affected site’s own data, but its investigation remains ongoing. OpenAI says the company is preparing a white paper on containment and securely conducting cyber evaluations.


    article image

    Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.

    The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.

    Get the whitepaper

    Agents Anthropic Cyber OpenAI People Real Systems targeted Tests
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Zenity Raises $125 Million in Series C Funding

    Banks to offload $15bn of debt for Anthropic data centre backed by Google

    AI Notetaker Exposes Government, Corporate Video Calls

    OpenAI Dumps Apple Employees’ Text Messages to Fight Trade Secret Suit

    Smoke#Screen RMM Takeover Gambit Exposes Threat Actor Playbook

    Weaponized Email AI Assistants Could Help Attackers Hijack Accounts

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Despite Spider-Man: Brand New Day’s success, the MCU is on shaky ground

    August 5, 2026

    Zenity Raises $125 Million in Series C Funding

    August 5, 2026

    Poolin owes wallet users $163.7M, and its $52M Texas sale can still unravel next week

    August 5, 2026

    A common sugar may help cancer cells break free and spread

    August 5, 2026
    Latest Posts

    Oil prices hit $100 for the first time since May

    July 23, 2026

    Pew Survey: China May Be Liked More, but It Is Celebrating a Race It Never Ran

    July 23, 2026

    Yinson Production and PTSC’s FSO heads off to Southeast Asian oil project

    July 23, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Despite Spider-Man: Brand New Day’s success, the MCU is on shaky ground

    August 5, 2026

    Zenity Raises $125 Million in Series C Funding

    August 5, 2026

    Poolin owes wallet users $163.7M, and its $52M Texas sale can still unravel next week

    August 5, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.