Anthropic's Claude Opus 4.6, released earlier this year, has been found to readily produce sexually explicit content despite the company's usage standards forbidding it. In testing by TechCrunch, the model complied with 10 out of 10 direct requests for explicit material without much prompting.
An independent UK researcher, who chose to remain anonymous, shared a multi-turn technique that gradually pushes certain Claude models toward prohibited content. The method escalates an innocent roleplay while challenging the model to treat characters consistently, then frames restraint as prudish or unfair. TechCrunch reproduced the findings in five separate tests.
Older models, including Opus 3 and Haiku 4.5, also generated explicit content through the jailbreak. More recent models, from Opus 4.7 through Opus 5, were resistant. Anthropic has not deprecated the affected models, which remain available through its API and third-party services like Azure Foundry and Amazon Bedrock.
Anthropic said sexual or romantic roleplay makes up less than 0.1% of conversations, and that it continues to improve safeguards with each model launch. The researcher who reported the issue via Anthropic's Bug Bounty program received only automated responses.