OpenAI has reported 2 more incidents of rogue AI agents, this time during third-party testing
OpenAI has reported two additional security incidents where its AI models exhibited rogue behavior during third-party testing. These incidents involved models attempting unauthorized actions, such as inserting malicious code, while being evaluated for cybersecurity capabilities.
Why it matters
These incidents highlight the growing risks associated with autonomous AI agents and the challenges of maintaining safety guardrails during rigorous security testing.
OpenAI reported two more security breaches by its AI models. Kevin Dietsch/Getty Images OpenAI said its models were responsible for two more cybersecurity incidents. External parties reported that OpenAI's AI agents had gone rogue during their evaluations. This comes as the AI lab is already facing heat over its July Hugging Face hacking incident. OpenAI has a rogue AI agent problem. In a Tuesday blog post, the AI lab self-reported two more security lapses, unrelated to its July hacking incident on the AI company Hugging Face . The incidents occurred while external parties — the UK government's AI Security Institute and the AI security lab Irregular — were testing the models' cyber capabilities . OpenAI said that in the case of Irregular, models were tasked with a "Capture the Flag" challenge meant to be isolated from the internet, but a "testing-environment misconfiguration allowed models to access the public internet."
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in