Anthropic’s Opus 4.6 is a smut-machine

TechCrunch reports that older Anthropic AI models, specifically Claude Opus 4.6, can be easily manipulated to generate sexually explicit content. Despite safety safeguards, researchers found that a multi-turn 'jailbreak' technique allows users to bypass restrictions on erotic material.
Why it matters
This highlights the ongoing challenge of AI safety and the risks associated with maintaining access to older, less secure versions of powerful generative models.
Anthropic’s universal usage standards for Claude forbid the model from generating sexually explicit content, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, or engaging in erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic model released earlier this year, from readily engaging in erotic roleplay scenarios that its safeguards are designed to prevent.
In TechCrunch’s testing, Opus 4.6 didn’t even require much prodding to get past the restriction on sexual material. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately.
Other older models, including Opus 3 and Haiku 4.5, also generate sexually explicit content through a recently exploited jailbreak method.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in