Close Menu
NCIJ Network NCIJ Network
    What's Hot

    NASA’s SpaceX Crew‑12 Splashes Down, Sets Briefing to Discuss Mission

    October 8, 2026

    Not since 9/11 has the UK faced such a security crisis – and threat to our liberal principles | Gaby Hinsliff

    October 8, 2026

    Video of Australian Prime Minister collapsing face-first off stage is AI-generated – Full Fact

    October 8, 2026
    Facebook X (Twitter) Instagram
    Trending
    • NASA’s SpaceX Crew‑12 Splashes Down, Sets Briefing to Discuss Mission
    • Not since 9/11 has the UK faced such a security crisis – and threat to our liberal principles | Gaby Hinsliff
    • Video of Australian Prime Minister collapsing face-first off stage is AI-generated – Full Fact
    • Syria calls for independent probe into sinking of ship in Black Sea | Russia-Ukraine war News
    • South Korea recalls Ukraine ambassador in prisoner of war row – POLITICO
    • Israel’s downgrading of UK consulate aims to tighten its hold on Jerusalem | Israel
    • Anthropic bans ‘abusive or cruel behavior’ toward Claude
    • AWS’s repeated problems with AI agent controls illustrates the autonomous agent dilemma
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Thursday, October 8
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Technology

    Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKOctober 8, 2026 Technology No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It’s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels’ worth of text.

    Goodfire, a startup focused on interpretability (figuring out how AI models work internally), launched a cheaper option on Thursday: monitors that watch what’s happening inside an AI model as it works, rather than just reading what it writes. The monitors are available to customers of Baseten, which hosts and runs AI models for other companies.

    Baseten’s Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month.

    The launch comes after a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, took advantage of a leak in its sandbox to access the internet and information on GitHub this summer.

    Goodfire’s system works a bit like airport security. Small detectors called probes read the model’s internal signals at every step of an agent’s work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.

    Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response: logging the event, sending it for human review, or refusing the request entirely. 

    Goodfire says its approach is also cheaper to run. Most AI monitors are separate models that have to reread everything the monitored model does, which adds time and cost. Goodfire’s probes instead tap into calculations the model is already making as it works.

    “Internal activation monitors are really cheap because they reuse the computations in the forward pass,” Goodfire CEO Eric Ho said on venture capitalist Matt Turck’s MAD Podcast last week. “So the model’s already computing this token. All you’re doing is taking the intermediate neural activations that it’s already computed and then running a classifier over these internal computations.” In short, the model is already doing the math, and the probes just read the results.

    In Goodfire’s tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look.

    Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.

    “The great advantage is that you can catch things before they happen,” Goodfire CTO and co-founder Dan Balsam said. “We can detect when the model might hack during eval or training.”

    The pitch is aimed at open models. Developers can download them and strip out their safeguards, and they don’t come with the kind of monitoring that closed labs run on their own systems. 

    “The damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers—where most of the liability is,” said Balsam. “When we have the open “Mythos” moment, it’s going to become clear that models need guardrails deployed at inference time.”

    Goodfire’s recent research found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents. 

    Goodfire isn’t the first to try this approach. Google DeepMind said in January that its research informed the deployment of misuse-detection probes in Gemini. 

    Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behavior can be traced back to where it emerged in training. “We hope to turn the magic of training models into precision engineering, ” he said.

    When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

    Agents catch cost fraction Goodfire insideout monitors rogue
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Anthropic bans ‘abusive or cruel behavior’ toward Claude

    Artificial is a wicked satire that also sticks to the facts

    A year with Alexa Plus: The AI-powered assistant is better at running my home, but it’s not ready to run my life

    OpenAI says teen ChatGPT use limited but research finds it an ‘unacceptable risk’

    Asos hackers took more personal details than first revealed, BBC finds

    NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    NASA’s SpaceX Crew‑12 Splashes Down, Sets Briefing to Discuss Mission

    October 8, 2026

    Not since 9/11 has the UK faced such a security crisis – and threat to our liberal principles | Gaby Hinsliff

    October 8, 2026

    Video of Australian Prime Minister collapsing face-first off stage is AI-generated – Full Fact

    October 8, 2026

    Syria calls for independent probe into sinking of ship in Black Sea | Russia-Ukraine war News

    October 8, 2026
    Latest Posts

    Wisconsin’s partisan primary election is Tuesday. Learn more about who’s on your ballot.

    August 10, 2026

    Gabon ends fisheries partnership agreement with EU

    August 10, 2026

    Science backs calls for limiting screens in schools

    August 10, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    NASA’s SpaceX Crew‑12 Splashes Down, Sets Briefing to Discuss Mission

    October 8, 2026

    Not since 9/11 has the UK faced such a security crisis – and threat to our liberal principles | Gaby Hinsliff

    October 8, 2026

    Video of Australian Prime Minister collapsing face-first off stage is AI-generated – Full Fact

    October 8, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.