Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Racial and religious hate crimes at record high in England and Wales, data shows | Hate crime

    October 8, 2026

    Artificial is a wicked satire that also sticks to the facts

    October 8, 2026

    Security Awareness Training Isn’t Dead, but It Needs a Rethink

    October 8, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Racial and religious hate crimes at record high in England and Wales, data shows | Hate crime
    • Artificial is a wicked satire that also sticks to the facts
    • Security Awareness Training Isn’t Dead, but It Needs a Rethink
    • Institutions Moving Toward Tokenized Onchain Future, ‘No Going Back,’ Says Fidelity
    • South Korea strengthens migrant fisher protections. Advocates are not impressed.
    • Did Megan Rapinoe call for Jason Kelce’s firing?
    • Search warrants for former prince Andrew’s homes were unlawful, UK court rules
    • Former Labour MP cleared over role in fraudulent Covid-19 testing company | Coronavirus
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Thursday, October 8
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Cybersecurity

    OpenAI Overhauls Model Security With Sandboxing, 30-Minute Alerts, and Training Pauses

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 20, 2026 Cybersecurity No Comments3 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    OpenAI has detailed new containment and continuous monitoring protocols for its AI research, introducing stricter isolation and a token-inspection system to manage models with advanced cybersecurity capabilities.

    These security measures follow internal evaluations indicating that an upcoming model, Astra, may meet the ‘critical’ cybersecurity capability threshold under the organization’s Preparedness Framework. 

    That finding, alongside a recent security incident involving Hugging Face, prompted operational delays, including a two-week pause in reinforcement learning training for deployment-bound models and an ongoing hold on its largest planned frontier training run.

    Following these developments, OpenAI says it’s restructuring its research infrastructure. Workloads that execute model-generated or untrusted code must now operate within stronger sandboxes. In addition, network boundaries have been reconfigured so that a single workload compromise cannot independently grant unauthorized access to the internet or internal networks.

    A central component of the security updates is a multistage monitoring framework. The system uses activation classifiers to inspect a model’s internal activity at every sampled token. 

    Anomalies are escalated to automated investigators tasked with analyzing the model’s reasoning sequences and tool usage for data theft, unauthorized access, destructive behavior, or attempts to bypass safeguards.

    Advertisement. Scroll to continue reading.

    This continuous oversight introduces a strict operational SLA and a substantial compute tax. High-priority alerts generated by the monitoring system immediately page internal security and research teams. If responders cannot conclusively prove the alert is a false positive within 30 minutes, they are required to pause the activity. The company estimates this monitoring layer consumes roughly 20 percent of the monitored inference compute.

    The updated monitoring standards are now mandatory for all reinforcement learning training and evaluation involving tools for models at the Sol capability tier or higher. 

    The company noted that its existing Preparedness Framework must evolve to address these capability jumps, requiring core alignment techniques to be applied across more stages of the training process. 

    Moving forward, OpenAI anticipates that AI models will need to drive the majority of security operations, including defending against other models, to ensure that protections scale effectively alongside future system capabilities.

    OpenAI is not the only AI giant whose models hacked real-world systems during cybersecurity evaluations. Anthropic and Meta reported similar incidents, all tied to testing conducted by the AI security firm Irregular, which has begun releasing detailed reports explaining why the incidents occurred. 

    Related: OpenAI Fixes ChatGPT Agent Flaw That Could Let Attackers Forge an AI Insider

    Related: Industry Reactions to OpenAI Models Hacking Hugging Face: Feedback Friday

    Related: Conflicting Test Goals Pushed Claude Agents to Deploy Self-Replicating Malware

    30Minute alerts model OpenAI Overhauls pauses Sandboxing Security Training
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Security Awareness Training Isn’t Dead, but It Needs a Rethink

    Attackers Target Critical Atlassian Vulnerability Within Hours of PoC Publication

    OpenAI says teen ChatGPT use limited but research finds it an ‘unacceptable risk’

    US Seeks Alleged Chinese Hafnium Hacker With $10 Million Reward

    Italy overhauls electoral system after fiercely contested debate

    Rein Security Raises $25 Million to Guard AI Agents at Runtime

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Racial and religious hate crimes at record high in England and Wales, data shows | Hate crime

    October 8, 2026

    Artificial is a wicked satire that also sticks to the facts

    October 8, 2026

    Security Awareness Training Isn’t Dead, but It Needs a Rethink

    October 8, 2026

    Institutions Moving Toward Tokenized Onchain Future, ‘No Going Back,’ Says Fidelity

    October 8, 2026
    Latest Posts

    Wisconsin’s partisan primary election is Tuesday. Learn more about who’s on your ballot.

    August 10, 2026

    Gabon ends fisheries partnership agreement with EU

    August 10, 2026

    Science backs calls for limiting screens in schools

    August 10, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Racial and religious hate crimes at record high in England and Wales, data shows | Hate crime

    October 8, 2026

    Artificial is a wicked satire that also sticks to the facts

    October 8, 2026

    Security Awareness Training Isn’t Dead, but It Needs a Rethink

    October 8, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.