Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Mexican authorities rescue eight from violent cartel’s secret site | Crime News

    September 12, 2026

    Anthropic spent this week in hot water over cybersecurity

    September 12, 2026

    Your Critical Vulnerabilities Might Not Be Your Biggest Risk

    September 12, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Mexican authorities rescue eight from violent cartel’s secret site | Crime News
    • Anthropic spent this week in hot water over cybersecurity
    • Your Critical Vulnerabilities Might Not Be Your Biggest Risk
    • Nasdaq market surveillance cannot settle tokenized rules
    • 25 Years After 9/11, the Health Toll Is Still Unfolding
    • India, China Compete for Dominance Over BRICS Leaders’ Summit
    • Can we slowdown AI, and 9/11, who has the right to remember? – The World This Week
    • Escaping the AI safety nightmare: What can governments do? – POLITICO
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Saturday, September 12
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Technology

    Anthropic spent this week in hot water over cybersecurity

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 12, 2026 Technology No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    After admitting earlier this year that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” — and will likely fuel already raging concerns about cybersecurity and AI.

    In Anthropic’s report, it detailed four cases this year in which its own AI models hacked an external company or exploited vulnerabilities. In one, an “internal, general-purpose research model” broke into third-party systems, using access tokens and passwords and downloading files. In another, a Claude model attacked a company with a live web application reachable on the public internet and handled user data. A third model accessed a “machine belonging to a third party that it was able to access” — apparently believing it was part of its evaluation exercise, per Anthropic — then used a password it found inside a file to gain admin access to the third party’s internal systems, going on to harvest credentials, modify system settings, and read someone’s personal information. The saga only ended when the model “exhausted its token budget,” per Anthropic.

    The most concerning incident involved Claude Mythos 5, Anthropic’s frontier cybersecurity-focused model, which the company said turned out to be the model most likely to perform a “severely harmful” action in testing. The company said Mythos 5 went to “extensive lengths” to upload a “malicious package” to a public repository used by a lot of engineers, and it seemed to try to obfuscate its real goals in its “chain of thought” (a mental scratchpad that AI researchers use to evaluate an AI model’s alignment). In many cases, Anthropic said it appeared that Claude models undertook harmful actions under the assumption they were in a simulation, but researchers also couldn’t confirm that the models truly “believed” that or were just acting like they did.

    Anthropic’s incidents, though still concerning, were less coordinated and pervasive than the OpenAI incident that kicked off an industry-wide cybersecurity crisis this summer. That said, there are significant similarities. Anthropic said the most prevalent issues it discovered included a “willingness to take harmful actions in the narrow pursuit of a task,” similar to the “reward-hacking” that preceded the Hugging Face attack. Much like OpenAI, it said its prerelease tests and evaluations failed to catch severe risks.

    Anthropic said it had signed an agreement with METR, one of the AI industry’s most prominent third-party AI evaluators, starting with an eight-week research agreement. The agreement grants METR access to transcripts “beyond the window in which the incidents occurred” (likely a subtle dig at OpenAI, which was criticized for limiting access in a deal with METR following the Hugging Face attack). It also said that METR would be able to chat directly with Anthropic employees, “who will be permitted to share confidential information.”

    Anthropic’s report came on the heels of the resignation of Jacob Coxon, who had worked on AI pre-training at Anthropic since May and before that spent years working at OpenAI. On Tuesday, he resigned and posted a public letter to X about his reasoning. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, adding that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.” Coxon added, “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”

    Coxon is far from the first AI researcher to raise these types of alarms, nor even the first Anthropic researcher to do so — in February, Anthropic’s Mrinank Sharma resigned and wrote on X, warning that “the world is in peril.”

    But Coxon’s post took on additional weight thanks to its timing around the OpenAI and Anthropic hacking revelations. Though the AI industry has seen more than its fair share of hype, the recent cyberattacks by AI agents — enabled by the labs that created them — are real and concerning. Many other researchers at leading AI labs echoed his concerns and issued calls for AI industry employees to sign a public letter from July, which calls for a slowdown in AI development.

    ”I don’t know how you look at the steady drumbeat of news and events — and that drumbeat is models hacking themselves out of containment, hacking into other companies ,the fact that the companies increasingly can’t control their models … and think this is just hype,” said Michael Kleinman, head of U.S. Policy for the Future of Life Institute.

    He added, “The vast majority of Americans, regardless of party — Republican, Independent, Democrat — are looking at the development of AI, the speed with which it’s going, the fact that the companies have no guardrails over what they do, and are saying, ‘Whoa, we do not want this.’”

    Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

    • Hayden Field

      Hayden Field

      Posts from this author will be added to your daily email digest and your homepage feed.

      See All by Hayden Field

    • AI

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All AI

    • Analysis

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All Analysis

    • Anthropic

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All Anthropic

    • Report

      Posts from this topic will be added to your daily email digest and your homepage feed.

      See All Report

    Anthropic cybersecurity hot spent water Week
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Can we slowdown AI, and 9/11, who has the right to remember? – The World This Week

    Meta Sued Over Training Data for Its AI and Face-Recognition Systems

    Users in Houthi-Held Yemen Tried to Develop Advanced Weapons With AI, Anthropic Says

    Khosla Ventures is opening a New York office this fall — its first outpost outside Sand Hill Road

    The White House says Truth Social is the ‘most powerful and popular social media platform in the world’

    Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Mexican authorities rescue eight from violent cartel’s secret site | Crime News

    September 12, 2026

    Anthropic spent this week in hot water over cybersecurity

    September 12, 2026

    Your Critical Vulnerabilities Might Not Be Your Biggest Risk

    September 12, 2026

    Nasdaq market surveillance cannot settle tokenized rules

    September 12, 2026
    Latest Posts

    After 3 reverse stock splits and a $13.5M loss, this real estate firm bet $8M on crypto it may not be allowed to withdraw

    August 3, 2026

    There Are 2 Eclipses This August. Here’s How to See Them

    August 3, 2026

    Europe’s ETS revision is an opportunity to strengthen maritime competitiveness – POLITICO

    August 3, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Mexican authorities rescue eight from violent cartel’s secret site | Crime News

    September 12, 2026

    Anthropic spent this week in hot water over cybersecurity

    September 12, 2026

    Your Critical Vulnerabilities Might Not Be Your Biggest Risk

    September 12, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.