Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI has developed an AI tool called GPT-Red to automate the red-teaming process and identify vulnerabilities in its models. This 'super-hacker' is designed to stay ahead of evolving cyber threats as AI agents become more complex.
Why it matters
As AI models gain more autonomy, automated safety testing is becoming a critical component of responsible AI development.
Exclusive: The firm says it wants to future-proof its safety procedures and stay ahead of human attackers.
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its most robust release yet.
GPT-Red automates a type of safety evaluation for software systems known as red-teaming, which is typically done by a team of human testers. The aim is to find as many different ways to break or hijack a system as possible. The weak spots can then be patched before the final version of the software is released.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in