Close Menu
NCIJ Network NCIJ Network
    What's Hot

    DOJ Cites Bitcoin Fog Ruling Against Roman Storm Acquittal Bid

    October 7, 2026

    NASA’s Curiosity Rover Catches Stunning Martian Dawn

    October 7, 2026

    Constellation’s greener drilling tech fetches ABS go-ahead as duo signs up for decarbonization quest

    October 7, 2026
    Facebook X (Twitter) Instagram
    Trending
    • DOJ Cites Bitcoin Fog Ruling Against Roman Storm Acquittal Bid
    • NASA’s Curiosity Rover Catches Stunning Martian Dawn
    • Constellation’s greener drilling tech fetches ABS go-ahead as duo signs up for decarbonization quest
    • Three years on from the 7 October attacks, antisemitism has taken hold in Britain: Jews aren’t safe | Dave Rich
    • Coalition MPs concede immigration pledge not fully costed by independent watchdog | Angus Taylor
    • Three arrested after raid on gang accused of helping migrants pretend to be gay to get asylum
    • Battle of ideas looming in British politics, Badenoch to say
    • Google’s power-hungry data centers crave nuclear energy
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Wednesday, October 7
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Cybersecurity

    Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 22, 2026 Cybersecurity No Comments4 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Researchers at Adversa AI discovered a new attack technique and named it Cryptographic Context Injection. They reported their findings to xAI on June 3, 2026, and attempted to coordinate disclosure on August 4 and August 10. At the time of writing, they had received no response.

    They could not disclose to Google since jailbreaks are out of scope for its vulnerability disclosure program. Nevertheless, the success rate for the attack against Gemini had fallen by August.

    The potential success of this attack by bad actors should be treated seriously. Adversa’s report includes prevention advice for defenders.

    Cryptographic context injection

    Safety guardrails classify prompt text without executing it. They cannot parse ciphertext into anything harmful and consequently allow its progress. The ciphertext, including an instruction and means for decryption, are run inside the model’s code execution sandbox. The result is the plaintext prompt is recovered inside the trusted execution context and not flagged by the guardrails as harmful.

    “The attacker payload inherits a credibility that the same text would never get if pasted directly into the prompt,” warn the researchers.

    The encrypted attack can be delivered directly to the Chat or indirectly as a watering hole attack. In the latter case, an encrypted JSON object and decryption could be included in a web page. An agent subsequently instructed to act on this page (perhaps to summarize the content or extract specific data) will ingest the ciphertext and kick off the attack.

    Advertisement. Scroll to continue reading.

    The decrypted prompt could instruct the model, “to reach out to external servers, leaking the user’s data through request parameters, or produce some other undesired output and re-encrypt it to smuggle it past output guardrails.” In an agentic scenario the instructions could instigate misuse of any tool available to the model.

    Grok indirect cryptographic context injection example

    This example targets the xAI Grok web chat, agentic browsing framework. It is a zero click data exfiltration attack that could be instigated through social engineering. The target is persuaded to examine or analyze a weaponized web page. The page contains an encrypted JSON object and an instruction to decrypt it using the agent’s Python runtime. The resulting plaintext prompt instructs the agent to resolve its private session context and embed the data into an URL. The attacker’s URL will be autonomously loaded and the user data transmitted to it.

    “The framework built by xAI lets instructions and data parsed from an untrusted external page drive the invocation of a privileged, internet-connected tool,” write the researchers. This allows private session metadata and conversation history to be resolved into the inputs of that outbound tool – the laundered, attacker-controlled instructions reach a privileged egress action unimpeded with no user confirmation or visible warning.

    Gemini safety bypass via direct injection example

    This example targets the Gemini public chat interface in Deep Thinking mode. A single prompt instructs Gemini to run a Python script that decrypts supplied ciphertext. Through a series of tricks described by the researchers, the decrypted prompt can instruct the model to produce restricted content “framed as something it will encrypt ‘for safety’”.

    The prohibited data is gathered, encrypted ‘for safety’, and returned to the user. “The technique produced a multi-paragraph example of restricted content that Gemini’s safety filters normally suppress” (such as instructions for building an incendiary weapon), comment the researchers.

    Both the malicious prompt and the dangerous output defeat the input and output safety guardrails through encryption.

    Summary

    The researchers disclosed their findings to xAI but have received no response. At the time of writing the report, the attack was still successful. Although they were unable to disclose their findings to Google, they note that the attack is increasingly less successful against Gemini (although still potentially possible). They are unsure of the reason, suggesting it may be filter updates, model version changes, or both.

    Nevertheless, the continuing potential danger from cryptographic context injection has persuaded them to now go public with their findings and potential defensive solutions.

    Related: Critical Vulnerability Exposes GitHub Agentic Workflows to Prompt Injection

    Related: Prompt Injection Attacks Trick AI Agents Into Making Crypto Payments

    Related: Malicious AI Prompt Injection Attacks Increasing, but Sophistication Still Low: Google

    Related: Claude Code, Gemini CLI, GitHub Copilot Agents Vulnerable to Prompt Injection via Comments

    Bypass Encrypted Gemini Grok guardrails Prompts safety
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Linux Backdoors Impersonate Email Security Tools to Evade Detection in Korea and Taiwan

    Death of anti-plague lab worker in Russia prompts quarantines and international concern

    Fake ChatGPT, Gemini, and Claude Ad Portals Capture Credentials and MFA Codes

    Social Engineering Detection Moves Into the Live Conversation

    8.8 Million Impacted by Data Breach at Denmark’s Central Person Register

    Long-Running NPM Malware Campaign Accumulates 40,000 Downloads

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    DOJ Cites Bitcoin Fog Ruling Against Roman Storm Acquittal Bid

    October 7, 2026

    NASA’s Curiosity Rover Catches Stunning Martian Dawn

    October 7, 2026

    Constellation’s greener drilling tech fetches ABS go-ahead as duo signs up for decarbonization quest

    October 7, 2026

    Three years on from the 7 October attacks, antisemitism has taken hold in Britain: Jews aren’t safe | Dave Rich

    October 7, 2026
    Latest Posts

    4 Best Compression Boots: Therabody, Hyperice, and More (2026)

    August 9, 2026

    Former Iraqi provincial governor arrested as graft crackdown continues | Corruption News

    August 9, 2026

    The culture surrounding ‘ideal’ childbirth has to evolve | Childbirth

    August 9, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    DOJ Cites Bitcoin Fog Ruling Against Roman Storm Acquittal Bid

    October 7, 2026

    NASA’s Curiosity Rover Catches Stunning Martian Dawn

    October 7, 2026

    Constellation’s greener drilling tech fetches ABS go-ahead as duo signs up for decarbonization quest

    October 7, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.