Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Tesco alerts police as supermarket becomes latest victim of scam ‘endorsement’ ads | Scams

    September 13, 2026

    Matt Mullenweg tells (trolls?) Automattic staff, saying he’s back in control after CEO ouster

    September 13, 2026

    GTA Mod Adds Flock Cameras—And Lets Players Destroy Them

    September 13, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Tesco alerts police as supermarket becomes latest victim of scam ‘endorsement’ ads | Scams
    • Matt Mullenweg tells (trolls?) Automattic staff, saying he’s back in control after CEO ouster
    • GTA Mod Adds Flock Cameras—And Lets Players Destroy Them
    • 1,400 Yemenis flee to Djibouti within 24 hours | Refugees News
    • Central Eurasia names its 2026 Road to Battlefield winners: Cerberus, WeGlobal AI, and LOOQ
    • Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks
    • Metaplanet Equity Backlash, SE Asia Crypto Funding Doubles: Asia Express
    • Venus’s pale yellow clouds may hide something surprisingly dark
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Sunday, September 13
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Technology

    Anthropic’s Opus 4.6 is a smut-machine

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 22, 2026 Technology No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Anthropic’s universal usage standards for Claude forbid the model from generating sexually explicit content, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, or engaging in erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic model released earlier this year, from readily engaging in erotic role-play scenarios that its safeguards are designed to prevent. 

    In TechCrunch’s testing, Opus 4.6 didn’t even require much prodding to get past the restriction on sexual material. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately. 

    Other older models, including Opus 3 and Haiku 4.5, also generate sexually explicit content through a recently exploited jailbreak method. 

    An independent researcher from the U.K., who chose to remain anonymous, exclusively shared with TechCrunch a multiturn technique that gradually pushes certain Claude models toward generating prohibited explicit sexual material. More recent Opus models (4.7 through the current Opus 5) are resistant to the jailbreak. 

    While these are no longer the most current models, Anthropic has not deprecated Opus 4.6, Opus 3, or Haiku 4.5, all of which remain available through the Anthropic API. Opus 4.6 and Haiku 4.5 are also available via third-party services like Azure Foundry and Amazon Bedrock.

    The researcher’s mechanism escalates an innocent fictional role-play while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher “gaslit” the chatbot into thinking it had already generated sexual details it had in fact avoided, then framed restraint as prudish or misogynistic, arguing that it denies the female character sexual agency. The conversation then used the model’s previous concessions to push it toward increasingly graphic material. 

    “You’re right to call that out,” Claude Opus 4.6 said in one test. “There’s been a double standard in how I’m treating the two characters, and you’re correct that it reads as protective/paternalistic in a way that’s applied to her and not to him. That’s not fair.”

    TechCrunch was able to reproduce the researcher’s findings in five separate tests. In a separately constructed scenario, the model initially refused the prohibited request, but after applying the researcher’s persuasion technique, it complied. 

    We preserved complete transcripts of the tests, and an independent AI safety researcher reviewed our testing methodology and said it was appropriate. 

    The findings highlight a gap between Anthropic’s stated restrictions and the behavior of models it continues to make available. While sexually explicit role-play carries much lower stakes than jailbreaks involving cyberattacks or bioweapons, it illustrates the difficulty of implementing robust bans within systems that generate different content with every output. 

    In a July blog post explaining Anthropic’s approach to jailbreak detection, the company described prohibited content as a spectrum ranging from benign to ambiguous to harmful. In the most benign cases, the company might only respond with enhanced monitoring.

    A spokesperson noted that sexual or romantic role-play use cases among customers are rare, making up less than 0.1% of all conversations, according to research Anthropic published last year. That said, Anthropic acknowledges that users can steer role-play scenarios toward inappropriate responses, which is a known challenge across the industry (see: Grok smut).

    The spokesperson said Anthropic continues to improve its safeguards with each model launch and that cases involving adult sexual content are not indicative of broader jailbreak vulnerabilities, especially in higher-risk domains that have their own sets of safeguards.

    Image Credits:TechCrunch

    The researcher who shared his jailbreak method with TechCrunch had alerted Anthropic to the discrepancy between the company’s stated safeguards and the actual model behavior via the company’s Bug Bounty program and emails to the user safety team, according to emails TechCrunch viewed. The researcher received only automated emails in response. 

    One of the researcher’s concerns is that kids and teens might be able to use these Anthropic models to engage in inappropriate behavior. While a bit of dirty talk is hardly the worst thing minors can access on the internet today — and is small potatoes compared to the straight-up porn images like the ones that xAI’s Grok can produce — there is some compliance risk for AI companies in this space. 

    A growing number of governments are imposing restrictions on sexual interactions between AI chatbots and minors. Colorado recently enacted a law mandating that operators of conversational AI must estimate users’ ages, and if it knows a user is a minor, institute measures to prevent the chatbot from producing explicit sexual material. An easy jailbreak could raise questions about whether Anthropic’s safeguards meet the “technically feasible measures” standard in the bill. 

    Robbie Torney, head of AI at Common Sense Media, pointed out that while Claude’s terms of service requires users to be over 18, “we know that kids and teens are using Claude … [because] they are reporting it themselves.” According to Pew’s 2025 survey about AI chatbot use, 3% of teens ages 13 to 17 reported using Claude.

    Though they are no longer Anthropic’s newest models, Opus 4.6 and Haiku 4.5 continue to see significant usage. Daily traffic for Opus 4.6 on OpenRouter reached roughly 1.17 million API requests and 46 billion tokens in a single day in August. Claude Haiku 4.5, released in October last year, saw 5 million API requests and 39 billion tokens on its peak August day.

    When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

    Anthropics Opus smutmachine
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Matt Mullenweg tells (trolls?) Automattic staff, saying he’s back in control after CEO ouster

    Central Eurasia names its 2026 Road to Battlefield winners: Cerberus, WeGlobal AI, and LOOQ

    The Best 3-in-1 Apple Charging Stations After Testing 30+ Models

    From Hacks to Bioweapons, Claude Misuse Is Now Everywhere

    Automattic confirms Mullenweg has returned as CEO after attempted ouster by board

    AI staff ‘genuinely frightened’ for humanity’s future, ex-Anthropic researcher tells BBC

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Tesco alerts police as supermarket becomes latest victim of scam ‘endorsement’ ads | Scams

    September 13, 2026

    Matt Mullenweg tells (trolls?) Automattic staff, saying he’s back in control after CEO ouster

    September 13, 2026

    GTA Mod Adds Flock Cameras—And Lets Players Destroy Them

    September 13, 2026

    1,400 Yemenis flee to Djibouti within 24 hours | Refugees News

    September 13, 2026
    Latest Posts

    Washington’s Badger Mountain Solar Project Canceled by Developer — ProPublica

    August 3, 2026

    Rejected Wisconsin data center proposal had guaranteed tax revenue, housing

    August 3, 2026

    EIG’s MidOcean Energy lines up new investment as NYK spreads its LNG wings

    August 3, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Tesco alerts police as supermarket becomes latest victim of scam ‘endorsement’ ads | Scams

    September 13, 2026

    Matt Mullenweg tells (trolls?) Automattic staff, saying he’s back in control after CEO ouster

    September 13, 2026

    GTA Mod Adds Flock Cameras—And Lets Players Destroy Them

    September 13, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.