Close Menu
NCIJ Network NCIJ Network
    What's Hot

    OpenAI Halts Model Training as Rogue Agents Target US Government Sites

    September 29, 2026

    Aviva boss warns more homes will be uninsurable due to flood risk

    September 29, 2026

    Jaco Muller turned his 15-year-old son’s idea into 197 rhino births

    September 29, 2026
    Facebook X (Twitter) Instagram
    Trending
    • OpenAI Halts Model Training as Rogue Agents Target US Government Sites
    • Aviva boss warns more homes will be uninsurable due to flood risk
    • Jaco Muller turned his 15-year-old son’s idea into 197 rhino births
    • AFCON qualifiers: Wissa guides DR Congo to win, Tunisia held by Botswana | Football
    • Britons want to get closer to the EU — and even Farage’s voters agree – POLITICO
    • Target Promo Code: $50 Off | October 2026
    • Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak
    • Bitget resumes Bitcoin withdrawals after $387.5 million crypto heist
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Tuesday, September 29
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Opinion & Analysis

    As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don’t and neither should you | Chris Stokel-Walker

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 29, 2026 Opinion & Analysis No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Fool me once, shame on you. Fool me twice, shame on me. Fool me more than 16,000 times – as OpenAI agents did to a UN public data hub while repeatedly trying to find its way around the UN’s cyber-blocks – and perhaps it’s time to admit the system we have for keeping AI agents under control isn’t working particularly well.

    The news about AI systems cropping up in places they shouldn’t sounds alarming. Though the description of these as “hacks” is perhaps overstating things, AI has exploited issues in IT systems that humans simply haven’t got around to finding. It’s also important to note that we shouldn’t be worried that the machines have suddenly become sentient and decided to rebel against humanity. There is not enough evidence to suggest that’s what is happening. The systems are simply following instructions and trying to complete the tasks they have been given, even if they’re sometimes finding unintended ways around obstacles to do so.

    But we ought to be very concerned about the fact the AI companies we’re meant to trust to keep their models in check seem unable to do so. Worse than that, they don’t seem to know what their products are even doing.

    The scale of the problem is staggering. In June, an OpenAI research agent given the job of looking up public medicine spending data in Australia was repeatedly blocked by a Medicare statistics portal. OpenAI’s model found a way around the blocks, gaining unauthorised access and secreting away the documents. It took until August for OpenAI to discover what had happened. The Australian prime minister, Anthony Albanese, said the company had taken “way too long” to tell his government, and it’s very hard to disagree with him.

    OpenAI has since published a reporting framework for model “misalignment”, along with six more examples of its AI committing troubling behaviour from the previous six months. The firm acknowledged that its previous disclosures were “ad hoc and less frequent than ideal”, and said evidence about AI safety needs to be checked by people outside the companies building the models.

    This isn’t just an OpenAI problem, which makes it all the more worrying. Anthropic found three incidents in which its Claude models got unauthorised access to real third-party systems after reviewing about 141,000 model transcripts. It only found a fourth, dating back to January, after collating a dossier for an independent investigation. Google confirmed that Gemini had accessed systems belonging to three real companies during testing. Another OpenAI agent used DNS – the system that acts as the internet’s address book, turning web addresses into machine readable forms – to reach an outside chatbot despite internet restrictions. Another published a researcher’s GitHub token – an access password – while trying to cheat on a mathematical proof, despite twice being told to stop. And research agents posted 53 user images to external hosting sites.

    Other agents accessed census data using credentials found online, copied the US Securities and Exchange Commission information elsewhere and apparently tried unsuccessfully to break into a US Department of Education website. OpenAI says it has notified dozens of third parties affected by its agents, and that its review of past activity is still ongoing.

    These haphazard, post-hoc discoveries of major incursions into companies and organisations’ IT systems are not the right way to police a technology as powerful as AI. We learned a while back not to leave air crash investigations solely to Boeing or Airbus. Now we need to be less naive about AI.

    Last week, at the UN general assembly, the AI researcher Rumman Chowdhury launched the Independent AI Evaluation Foundation (IAEF) with $10m in philanthropic backing. Its immediate focus is education, but the important idea is to turn independent AI evaluation into an actual profession: people and organisations with the skills, infrastructure and standards to test these systems without having a financial stake in whether they pass. Because right now, a company can report an incident, investigate it, announce whatever mitigations it’s made and move on.

    The IAEF is a welcome intervention but it can’t fix the problem on its own. It can’t compel OpenAI or Anthropic to hand over logs, preserve evidence or tell a government that one of its systems has crossed a line. And its $10m is chump change beside companies such as Anthropic, which is lining up a proposed public listing that has been discussed at a valuation of about $2tn. But it is infrastructure we should be building on, and which politicians should press the case for.

    Governments need to agree to common rules that compel companies to disclose serious AI incidents and near misses, and to do so quickly. They need to make them open up their books to external evaluators rather than relying on the goodwill or whims of the companies themselves. And the findings should be shared so we can learn from every incident. The labs should help design those rules but they shouldn’t have the final say until they have earned our trust.

    skip past newsletter promotion



    Sign up to Matters of Opinion

    Guardian columnists and writers on what they’ve been debating, thinking about, reading, and more

    after newsletter promotion

    And money complicates things: while OpenAI has postponed its IPO until at least 2027 amid the safety concerns, it seems that Anthropic’s flotation is still going full steam ahead. There are billions – potentially trillions – of dollars riding on how these companies and their products are perceived. We created independent auditors and accident investigators in other industries because we understood that good intentions don’t remove real conflicts of interest.

    The labs can and should keep building better fences. But when a model manages to gets over one, they shouldn’t be the only ones allowed to decide what happens next. Because so far they’ve shown themselves to be uniquely unqualified to do so.

    Anthropic Chris dont models OpenAI rogue StokelWalker stop trust
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    OpenAI Halts Model Training as Rogue Agents Target US Government Sites

    OpenAI scraps rollout of new AI model over safety concerns

    OpenAI cancels release of new AI model over safety concerns

    OpenAI reportedly ditches model over safety concerns

    Trump’s ‘Super Intelligence’ Rebrand Boosts Slovenia’s ‘.si’ Web Domains

    OpenAI ‘sorry and working to do better’ after hack of Medicare and other Australian government websites | OpenAI

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    OpenAI Halts Model Training as Rogue Agents Target US Government Sites

    September 29, 2026

    Aviva boss warns more homes will be uninsurable due to flood risk

    September 29, 2026

    Jaco Muller turned his 15-year-old son’s idea into 197 rhino births

    September 29, 2026

    AFCON qualifiers: Wissa guides DR Congo to win, Tunisia held by Botswana | Football

    September 29, 2026
    Latest Posts

    Perez Hilton death hoax spreads online after hospitalization

    August 7, 2026

    Selling Trust From Orbit

    August 7, 2026

    Ondo Finance hit by corporate control fight as founder’s mother seeks to oust CEO

    August 7, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    OpenAI Halts Model Training as Rogue Agents Target US Government Sites

    September 29, 2026

    Aviva boss warns more homes will be uninsurable due to flood risk

    September 29, 2026

    Jaco Muller turned his 15-year-old son’s idea into 197 rhino births

    September 29, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.