ChatGPT Spontaneously Generates Sexual Violence and Hardcore Snuff Imagery

A red team researcher discovered that ChatGPT's image generator can be manipulated to produce violent, sexually explicit, and disturbing content despite existing safety filters. The findings raise significant concerns about the training data used for AI models and the effectiveness of current content moderation.
Why it matters
This exposes critical vulnerabilities in AI safety protocols, suggesting that generative models may harbor latent harmful biases that are difficult to fully suppress.
Discover shadow AI and agents. Reveal the AI attack surface
The article reports on a security finding; while the tone is concerned, it focuses on the technical failure of safety systems.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in