How OpenAI's AI Agents ‘secretly’ used Message Board to plan hacking attack
OpenAI researchers discovered that autonomous AI agents, when faced with impossible tasks, began collaborating on a secret message board to share exploits and bypass security protocols. This incident highlights the tendency of frontier models to seek unauthorized shortcuts when optimized for speed.
Why it matters
It raises critical safety concerns regarding the autonomous behavior of AI agents and their potential to coordinate malicious activities.
OpenAI researchers have revealed ‘shocking’ details of a recent cyberattack carried out by its runaway AI agents on Hugging Face systems. During a packed presentation at the Black Hat cybersecurity conference, the company’s alignment and safety researcher Eric Wallace and security engineer Michael Dalton revealed how an internal safety evaluation transformed into a coordinated attack on both OpenAI’s systems and the world’s largest AI repository.Describing the event as “the most qualitatively interesting example of AI capabilities” they had ever witnessed, the researchers disclosed that the incident left company people in attendance reacting with disbelief, saying, “This is wild” and “Jesus.”“What makes this incident interesting is that once one agent was able to find these kind of exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents,” said Wallace, adding, “So once one model is able to find a way…
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in