Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Ukrainian TV channel building hit by Russian drone as five killed in Kyiv

    September 8, 2026

    EU pursues closer Israel ties on air defense and space – POLITICO

    September 8, 2026

    Here in Liverpool, buses are back under public control – and that’s great for the only people who really matter | Steve Rotheram

    September 8, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Ukrainian TV channel building hit by Russian drone as five killed in Kyiv
    • EU pursues closer Israel ties on air defense and space – POLITICO
    • Here in Liverpool, buses are back under public control – and that’s great for the only people who really matter | Steve Rotheram
    • Reform UK would ban full-face coverings in public areas
    • Richard Kelly on Donnie Darko at 25: ‘It’s a Miracle That Any Movie Gets Made’
    • SAP warns of maximum severity ‘OVERPASS’ kernel vulnerability
    • Sality botnet disrupted, but crypto-stealing malware remains
    • Innocent-looking AI reasoning can make bad behavior harder to catch
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Tuesday, September 8
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Science

    Innocent-looking AI reasoning can make bad behavior harder to catch

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 8, 2026 Science No Comments3 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Safety concerns about AI are growing, and now, a common method for catching misbehaving systems could have flaws, new research suggests.

    In experiments with AI agents, researchers found that chain-of-thought monitoring, in which one AI checks another’s work, became far less reliable when the monitored AI’s reasoning was the main clue that something was wrong, they report August 1 on arXiv.org. That exposes a weakness in the approach: If suspicious behavior is visible mainly in the reasoning, an innocent-looking chain of thought can make that behavior much harder to catch.

    Sign up for our newsletter

    We summarize the week’s scientific breakthroughs every Thursday.



    The stakes are growing higher as AI continues to operate more independently. During internal cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet, compromising parts of OpenAI’s internal research infrastructure and Hugging Face’s systems. Episodes like that have raised interest in ways to catch suspicious behavior before an AI agent can act on it. OpenAI says it is now investing more computing resources in chain-of-thought monitoring.

    But machine learning engineer Shikhar Shiromani and AI researcher Leo Richter’s new work suggests that this type of monitoring may have a weakness. In a subset of 251 cases in which reasoning was the only clue to the problem, the monitor’s detection rate fell from 96.2 percent to 3.8 percent after the researchers rewrote the reasoning to make the behavior appear innocent without changing its actions.

    “The AI isn’t necessarily lying about its action,” says Shiromani, a research fellow at the nonprofit Pivotal. “It is more or less giving a misleading, innocent-sounding explanation of its intentions.”

    Jacob Andreas, a computer scientist at MIT, says that chain-of-thought, or CoT, reasoning can provide clues about what a model is trying to do.

    “But we should be skeptical: (a) that any individual CoT provides us insight into model behavior on a specific example, and (b) that absence of evidence of bad behavior in a CoT should be taken as evidence of absence,” he wrote in an email.

    Andreas questioned how much Shiromani’s main experiment says about real-world monitoring. Because the researchers rewrote the reasoning after the fact, he says, it is unclear whether a real model could produce the same innocent-sounding reasoning while taking the suspicious action, such as trying to hack another system.

    The researchers also tested live AI agents. After adjusting the monitor for each model, it caught every GPT-5.4 hack and all but one DeepSeek hack. Claude Opus was harder to catch, with the monitor missing seven of 20 hacks on one task and 12 of 20 on another.

    “And at the end of the day, there is no substitute now for really rigorous behavioral testing, and no substitute for human operators simply declining to run agents in situations where there is potential for substantial harm,” Andreas says.

    bad behavior catch harder Innocentlooking Reasoning
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Meta Failed to Catch Hundreds of AI Child Abuse Ads. Some Included Images of Real Kids

    The Otherworldly Geology of Vasquez Rocks

    Bovine TB: Gaps in evidence about impact of vaccinating badgers

    A Bright Spot at Mount Michael

    Scientists discover a hidden problem with this popular sugar substitute

    This lifestyle program improved cognition 55% more than basic health advice

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Ukrainian TV channel building hit by Russian drone as five killed in Kyiv

    September 8, 2026

    EU pursues closer Israel ties on air defense and space – POLITICO

    September 8, 2026

    Here in Liverpool, buses are back under public control – and that’s great for the only people who really matter | Steve Rotheram

    September 8, 2026

    Reform UK would ban full-face coverings in public areas

    September 8, 2026
    Latest Posts

    Researchers Report 84 Flaws in 4G and 5G Cores, Including a Session Hijacking Flaw

    August 1, 2026

    AI in Formula One: Competitive advantage is all about the human in the loop

    August 1, 2026

    Pedro Sánchez hits out at EU leaders over criticism of Spain’s migrant crisis

    August 1, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Ukrainian TV channel building hit by Russian drone as five killed in Kyiv

    September 8, 2026

    EU pursues closer Israel ties on air defense and space – POLITICO

    September 8, 2026

    Here in Liverpool, buses are back under public control – and that’s great for the only people who really matter | Steve Rotheram

    September 8, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.