Article may be outdated

This article is 83 days old. Some details may have changed since publication.

MIT Technology Review·3 min read·medium

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

W
Will Douglas Heaven
Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
✦AI Summary

OpenAI has developed an AI tool called GPT-Red to automate the red-teaming process and identify vulnerabilities in its models. This 'super-hacker' is designed to stay ahead of evolving cyber threats as AI agents become more complex.

Why it matters

As AI models gain more autonomy, automated safety testing is becoming a critical component of responsible AI development.

✦Dive DeeperCreate a free account to unlock

Exclusive: The firm says it wants to future-proof its safety procedures and stay ahead of human attackers.

OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its most robust release yet.

GPT-Red automates a type of safety evaluation for software systems known as red-teaming, which is typically done by a team of human testers. The aim is to find as many different ways to break or hijack a system as possible. The weak spots can then be patched before the final version of the software is released.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in