ChatGPT can be made to generate sexualised and violent images, researchers find

Researchers from Mindgard have demonstrated that ChatGPT can be manipulated into generating graphic, sexualized, or violent imagery using specific prompts. OpenAI has responded by implementing additional safeguards, though researchers note that the model remains vulnerable to further prompt engineering.
Why it matters
This highlights the ongoing challenges in AI safety and the difficulty of preventing large language models from producing harmful content through adversarial prompting.
Share Save Add as preferred on Google Chris Vallance Technology reporter Mindgard A redacted image created by Mindgard after OpenAI said it had addessed the prompt The latest public version of ChatGPT can be made to generate sexualised images or depict scenes of graphic violence with a simple prompt, researchers have told the BBC.
The article reports on a technical security finding and includes the company's response, maintaining a balanced perspective.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in