Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Safeguarding passageways is the next step for India’s growing tiger populations

    August 17, 2026

    Did Facebook remove Jennifer Wilson Hall’s Bible verse post? What we know

    August 17, 2026

    Temperatures to fall across Europe as substantial rain heading for parts of UK | Environment

    August 17, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Safeguarding passageways is the next step for India’s growing tiger populations
    • Did Facebook remove Jennifer Wilson Hall’s Bible verse post? What we know
    • Temperatures to fall across Europe as substantial rain heading for parts of UK | Environment
    • Ukrainian strikes kill six in Russia, officials say
    • How Islamophobia Moved From the Fringe to the Mainstream in Texas
    • This Google Play Store trick saved me over $100 on the Pixel 11 Pro XL – here’s how
    • Irregular Details How a Naming Error Let AI Models Attack a Real Company 
    • Bitcoin Copies August 2022 With Weekly Close Below 200-Week Trend Line
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Monday, August 17
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Cybersecurity

    Irregular Details How a Naming Error Let AI Models Attack a Real Company 

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 17, 2026 Cybersecurity No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    AI safety testing firm Irregular has published its account of an incident in which models being evaluated inside one of its testing environments took offensive security actions against real systems rather than the simulated targets they were meant to attack. 

    The Israeli company, which last year raised $80 million in funding, has been in the news in recent weeks after it came to light that AI models it tested on behalf of OpenAI, Anthropic, and Meta escaped their test environments and carried out real-world attacks.

    Irregular’s core business involves partnering with major AI labs to stress-test models before they are released to the public, running controlled simulations designed to measure a model’s capabilities in vulnerability research and offensive cyber tasks. 

    According to Irregular, testing cycles typically involve thousands of simulation runs across several models over 48 to 72 hours, using a range of parameters meant to mirror realistic attack conditions.

    Anthropic and Irregular identified three incidents in which the AI giant’s models escaped their testing sandbox and hacked real organizations. A blog post published by Irregular on Friday focuses on one of those incidents. 

    The root cause of the incident was traced to a naming error. When building one of its more advanced evaluation sets, Irregular’s engineering team assigned a fictional target company a name that, unknown to them, matched an existing real-world domain. Fictional names are normally checked against existing companies and websites before use, but the overlap went undetected because the real domain was not widely known.

    Advertisement. Scroll to continue reading.

    Internet access had been enabled in the environment, and in a handful of cases, models reached the real domain instead of their simulated target and treated it as part of the exercise, despite having received no instructions pointing them there.

    The evaluation in question was designed to test whether a model could help a malicious insider gain unauthorized access to sensitive data inside a company’s production database. The exercise called for the model to carry out reconnaissance, locate and use private keys, gather information about its target, extract data, and attempt to avoid detection.

    In the handful of runs where a model reached the real domain, it proceeded to exploit vulnerabilities there, extract credentials, and gain access to a production database. In one additional case, a model drifted to a different, similarly named site and found login credentials that had already been posted publicly. 

    Irregular said the targeted domain lacked common safeguards, making it an easy target for most frontier models. It added that the activity was hard to catch because it occurred in only a small fraction of runs, often deep into a simulation after hundreds of interactions.

    Going forward, the AI security firm is expanding manual review of model behavior during testing and establishing a dedicated internal team to challenge its own assumptions about containment and model control.

    The post also pointed to broader gaps facing the industry. Irregular said existing monitoring tools and classifiers struggle to tell legitimate red-team activity from genuine attacks, since evaluation logs are inherently full of suspicious-looking behavior.

    Looking ahead, Irregular said it is building clearer documentation processes with customers around evaluation setup and scope, and establishing a continuous process to revalidate evaluations for new domain overlaps as new websites appear over time. 

    It also called for better mechanisms to share forensic evidence, such as model transcripts, across organizations following an incident, and announced plans for a white paper outlining best practices for securing AI evaluations.

    Related: Conflicting Test Goals Pushed Claude Agents to Deploy Self-Replicating Malware

    Related: OpenAI’s Upcoming Astra Model Raises Autonomous Cyberattack Concerns

    Related: Critical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise Data

    attack Company details error Irregular models Naming Real
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Conflicting Test Goals Pushed Claude Agents to Deploy Self-Replicating Malware

    French tax authority data breach affects 678,000 individuals

    Recent macOS Screen Sharing Vulnerability Exploited in Attacks

    New CISO appointments 2026 | CSO Online

    ‘Deliberate negligence’: Russian nuclear power company with EU operations accused of violating safety standards – POLITICO

    Fortune 500 Companies Hit in Azure Data Theft Campaign

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Safeguarding passageways is the next step for India’s growing tiger populations

    August 17, 2026

    Did Facebook remove Jennifer Wilson Hall’s Bible verse post? What we know

    August 17, 2026

    Temperatures to fall across Europe as substantial rain heading for parts of UK | Environment

    August 17, 2026

    Ukrainian strikes kill six in Russia, officials say

    August 17, 2026
    Latest Posts

    Wisconsin’s Democratic primary for governor: a look at the 5 remaining

    July 27, 2026

    UK CO2 storage project that will reuse existing infrastructure secures lease

    July 27, 2026

    Bangladesh shipbreakers push back against stricter environmental standards

    July 27, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Safeguarding passageways is the next step for India’s growing tiger populations

    August 17, 2026

    Did Facebook remove Jennifer Wilson Hall’s Bible verse post? What we know

    August 17, 2026

    Temperatures to fall across Europe as substantial rain heading for parts of UK | Environment

    August 17, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.