Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Green party votes to adopt policy that equates Zionism with racism | Green party

    October 4, 2026

    The Best Gifts Under $25 for Everyone on Your List (2026)

    October 4, 2026

    Anthropic asks Claude users to share voice data for AI model training

    October 4, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Green party votes to adopt policy that equates Zionism with racism | Green party
    • The Best Gifts Under $25 for Everyone on Your List (2026)
    • Anthropic asks Claude users to share voice data for AI model training
    • Could a stalled Clarity Act slow crypto dealmaking?
    • A meteor hit Oklahoma 100 million years later than scientists thought
    • The week in pictures: Disaster averted on Flydubai flight, French school protests and Bangkok floods
    • Moscow says it will step up strikes on Ukraine after Kyiv vows to hit Russian oil refineries – POLITICO
    • Jacob Rees-Mogg accuses Tucker Carlson of ‘verging on defending Hitler’ | Tucker Carlson
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Sunday, October 4
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Crypto & Blockchain

    ‘Inner Thoughts’ of Every Major AI Model Exposed in Massive Exploit

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 13, 2026 Crypto & Blockchain No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    In brief

    • A team or researchers found that Anthropic, OpenAI, and Google all use a single global encryption key for AI reasoning tokens.
    • By decoding 315,320 reasoning blocks scraped from public GitHub and Hugging Face repositories, the researchers recovered 182 credentials, including 62 live API keys, 33 passwords, and 30 personal email addresses.
    • OpenAI, Anthropic, and Google deployed server-side patches after responsible disclosure, but historical session logs already shared publicly remain decodable.

    Security researchers have found a way to read the encrypted “inner thoughts” of every major AI reasoning model—and uncovered 62 live API keys and 33 passwords buried in session logs that developers had shared publicly online without knowing what was inside them.

    “By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials,” the researchers wrote.

    The paper, submitted August 10 by a team from MATS Research, the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, and security firm Snyk, targets a specific class of AI: reasoning models. These are models that don’t just answer immediately and instead start with an internal chain-of-thought (a step-by-step scratchpad where the AI works through a problem before showing you the answer), then deliver a final response.

    Anthropic, OpenAI, and Google all encrypt that hidden scratchpad. Encryption—the process of scrambling data into an unreadable code—is meant to protect the company’s intellectual property and keep sensitive intermediate reasoning away from users. The encrypted block gets passed back to the provider’s servers with every follow-up message, maintaining the conversation without storing anything on the company’s end.

    One key to rule them all

    The flaw is architectural. Instead of binding each encrypted reasoning block to a specific user, session, or model, all three providers use a single, provider-wide encryption key across their entire ecosystem. “These encrypted blocks are fully compatible and interchangeable across different sessions, users, and even different models within a provider’s ecosystem,” the researchers wrote.

    That means a block of encrypted reasoning from Claude Opus 4.8—Anthropic’s flagship model—can be injected into Claude Haiku 4.5, a cheaper, less guarded sibling without breaking Anthropic’s rules. Haiku lacks the anti-distillation alignment (safety training specifically designed to stop a model from transcribing its own reasoning on command) that Opus has.

    Tell Haiku to read out the encrypted block verbatim, and it does. “By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly,” the paper states.

    “Cross-model portability means Haiku 4.5 can read Opus 4.8’s thoughts,” lead researcher Alexander Panfilov wrote on X. The same attack reproduced across OpenAI’s GPT-5.6 family and Google’s Gemini model lineup. No special access required—standard API access (the connection developers use to build applications on top of AI models) was sufficient to execute it.

    We can finally talk about it:

    We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.

    We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried. pic.twitter.com/S7wN8aP3X7

    — Alexander Panfilov (@kotekjedi_ml) August 11, 2026

    What the public logs contained

    To demonstrate real-world damage, the team scraped 6,708 publicly shared AI agent transcripts—automated session logs that developers routinely post to GitHub and Hugging Face for collaboration or debugging. They decoded 315,320 reasoning blocks from those logs.

    “Developers frequently share their session logs and encrypted thinking traces publicly online, entirely unaware of the sensitive data hidden within the encrypted blocks,” the paper notes. Most of those secrets never appeared in the visible AI output—they existed only inside the encrypted reasoning, invisible to anyone who hadn’t run the attack.

    The vulnerability opens four attack vectors beyond simple credential theft: stealing proprietary reasoning patterns from AI companies to train competing models via distillation (when a smaller AI learns to mimic a bigger one by studying its outputs); extracting private data from shared logs; executing invisible prompt injection, where malicious instructions are hidden inside encrypted reasoning blocks that security monitoring tools never see; and jailbreaking powerful models through their less-guarded siblings.

    Anthropic, OpenAI, and Google all deployed server-side mitigations after the team followed responsible disclosure procedures. As Decrypt previously reported, Anthropic has been a recurring focus for security researchers this year, especially as its latest models consume a lot more tokens in that process.

    The patches are live. The 6,708 session transcripts with decoded reasoning blocks already scraped from the public web are not going anywhere.

    Daily Debrief Newsletter

    Start every day with the top news stories right now, plus original features, a podcast, videos and more.

    exploit exposed major massive model Thoughts
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Anthropic asks Claude users to share voice data for AI model training

    Could a stalled Clarity Act slow crypto dealmaking?

    ‘We Have Identified You, Sir’: Near Intents Recovers $3.8 Million After 48-Hour Ultimatum

    Binance will hold some Brazil crypto deposits until users explain where the money came from

    The Pope Has Thoughts on AI Art—And They’re Not Flattering

    El Salvador gets $138 million from IMF

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Green party votes to adopt policy that equates Zionism with racism | Green party

    October 4, 2026

    The Best Gifts Under $25 for Everyone on Your List (2026)

    October 4, 2026

    Anthropic asks Claude users to share voice data for AI model training

    October 4, 2026

    Could a stalled Clarity Act slow crypto dealmaking?

    October 4, 2026
    Latest Posts

    Bald Range Wildfire Forces Evacuation of 18,000 in British Columbia

    August 9, 2026

    Amazon deforestation alerts fall to lowest level since 2013, Brazilian data show

    August 9, 2026

    Institutional bear market: Why Bitcoin’s downturn is different

    August 9, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Green party votes to adopt policy that equates Zionism with racism | Green party

    October 4, 2026

    The Best Gifts Under $25 for Everyone on Your List (2026)

    October 4, 2026

    Anthropic asks Claude users to share voice data for AI model training

    October 4, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.