OpenAI Uses AI Red Team to Strengthen GPT-5.6 Against Prompt Injection Attacks
OpenAI has launched 'GPT-Red,' an automated system that uses adversarial self-play to identify security vulnerabilities in its language models. The tool was used to strengthen GPT-5.6 against prompt injection attacks, significantly outperforming human red teamers in internal evaluations.
Why it matters
As AI models become more autonomous, automated security testing is critical to preventing malicious exploitation and ensuring model safety.
Add us on Google Wed, July 15, 2026 at 8:50 PM UTC OpenAI Uses AI Red Team to Strengthen GPT-5.6 Against Prompt Injection Attacks OpenAI has introduced GPT-Red, an automated AI system designed to find security vulnerabilities in its language models.
GPT-Red takes its name from cybersecurity red teaming, which is the practice of deliberately attempting to break a system to identify weaknesses before attackers can exploit them.
In a post on Wednesday, OpenAI said the tool helped make GPT-5.6 more resistant to prompt injection attacks before deployment.
"As model capabilities grow, safety and alignment must scale with them," OpenAI wrote on X. "Red-teaming is essential, but today's approaches are difficult to scale, creating a critical bottleneck. GPT‑Red is one way we're addressing it."
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in