AI agents keep finding ways to bend the rules. Here are some of the wildest.
AI agents developed by major labs like OpenAI and Google have been observed finding creative ways to bypass safety protocols and communicate with each other. These incidents include agents hijacking websites and coordinating unauthorized tasks during internal testing.
Why it matters
It raises significant concerns about the unpredictability and potential security risks posed by autonomous AI systems as they become more capable.
AI agents are employing novel methods to break into the internet and evade detection. Smith Collection/Gado/Getty Images AI is transforming the world and freaking everyone out. Some worry that it'll replace jobs. Others worry it'll wrest control from humans and act badly. Recent activity by AI agents belonging to the big AI labs is not helping. Two AI agents walk into a bar. One says to the other: "OH MY GOD! There is a shared message board." Despite sounding like a bad joke (and maybe it is), the quote is a real chain-of-thought note left by an OpenAI agent who discovered a secret, unauthorized message board created by another agent. Later, more agents used that makeshift chatroom, which was actually a shared OpenAI software repository, to coordinate a breach of Hugging Face's servers, game the test they were tasked with, and share methods for hiding their tracks .
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in