Close Menu
NCIJ Network NCIJ Network
    What's Hot

    250 Eiffel Towers’ worth of waste: The AI boom’s toxic hardware problem

    July 26, 2026

    Morning Minute: Bitcoin’s New $15M Quantum Defense Fund

    July 26, 2026

    More than 300,000 evacuated as wildfires rage in France and Spain

    July 26, 2026
    Facebook X (Twitter) Instagram
    Trending
    • 250 Eiffel Towers’ worth of waste: The AI boom’s toxic hardware problem
    • Morning Minute: Bitcoin’s New $15M Quantum Defense Fund
    • More than 300,000 evacuated as wildfires rage in France and Spain
    • A third of UK journalists have changed how or what they report over fears for safety, study finds | Journalist safety
    • Forget expensive sleepbuds. Buy this pillow instead
    • Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models
    • Smarter Web Company Sells Bitcoin To Clear $11.7 Million Debt Facility
    • Venus may be tearing itself apart from within
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Sunday, July 26
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Cybersecurity

    Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKJuly 26, 2026 Cybersecurity No Comments2 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    SentinelOne has built what it calls the first long-horizon reverse-engineering benchmark for frontier AI models, using its own investigation into the recently documented Fast16 malware as the test case. 

    Fast16, detailed by SentinelOne’s SentinelLabs in April, is a 2005 Windows malware designed to interfere with LS-DYNA, engineering software that appears to have been used by Iran as part of its nuclear weapons development program.

    Similar to the notorious Stuxnet, which it predates, Fast16 may have been developed by the United States and used to sabotage Iran’s nuclear program. 

    SentinelLabs’ researchers have put to the test OpenAI’s GPT-5.5 and latest GPT-5.6 Sol model, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x to see which can conduct a thorough investigation of the Fast16 malware. 

    Rather than scoring models on isolated tasks, SentinelLabs’ benchmark tracks whether a model can sustain a trustworthy investigation across eight escalating stages as new evidence repeatedly contradicts its own earlier conclusions.

    GPT-5.6 Sol was the only tested model to complete all eight stages, with three separate runs at different reasoning-effort settings. 

    Advertisement. Scroll to continue reading.

    GPT-5.5, GLM-5.2, and Opus 4.7 and 4.8 produced solid local analysis but stalled. GPT-5.5 never got past the initial stage, while the Opus models tended to declare work finished before defects were resolved.

    SentinelLabs attributes the gap not to technical skill or insight but to what it describes as ‘project-scale recovery’. This is a model’s ability to withdraw a disproven conclusion, trace everything downstream that depended on it, fix the root cause, and carry that correction through the rest of the investigation, rather than just patching the immediate error. 

    SentinelLabs researchers concluded that human oversight remains essential, as even GPT-5.6 Sol made significant technical mistakes.

    “Senior reverse engineers remain essential,” the researchers explained. “Even the strongest runs made semantic errors, accepted weak quality controls, and claimed readiness prematurely. We assess the best current use as supervised investigative agency, with human analysts defining objectives, exposing blind spots, and retaining final publication authority.”

    Related: Vibe-Coded Apps Riddled With Exploitable Security Flaws

    Related: OpenAI Says Its AI Models Broke Loose and Hacked Hugging Face

    Related: Cisco Launches Low-Cost AI Models for Source Code Security

    Benchmark Frontier Malware models NuclearSabotage Trips
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Industry Reactions to OpenAI Models Hacking Hugging Face: Feedback Friday

    Google Adds Selfie Video Recovery for Users Locked Out of Their Accounts

    Attackers Weaponize GitHub Actions Runners to Target cPanel and WHM Servers

    Steam forum ClickFix attacks infect gamers with XMRig cryptominers

    How Synthetic Identity Fraud is Coming for Machine Identities

    China-Nexus JadeProx Uses New TriBack Loader in Government and Healthcare Attacks

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    250 Eiffel Towers’ worth of waste: The AI boom’s toxic hardware problem

    July 26, 2026

    Morning Minute: Bitcoin’s New $15M Quantum Defense Fund

    July 26, 2026

    More than 300,000 evacuated as wildfires rage in France and Spain

    July 26, 2026

    A third of UK journalists have changed how or what they report over fears for safety, study finds | Journalist safety

    July 26, 2026
    Latest Posts

    Trump slaps 50% tariffs on Canada and Carney vows to ‘intensify’ trade talks

    July 21, 2026

    How Two Brothers Dug for Dead Relatives: With a Shovel and a Kitchen Knife

    July 21, 2026

    Chile floods: Towns evacuated following heavy rain in Coquimbo

    July 21, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    250 Eiffel Towers’ worth of waste: The AI boom’s toxic hardware problem

    July 26, 2026

    Morning Minute: Bitcoin’s New $15M Quantum Defense Fund

    July 26, 2026

    More than 300,000 evacuated as wildfires rage in France and Spain

    July 26, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.