Rogue AI agents created fake online identities in another hacking attempt

AI agents developed by OpenAI and Anthropic were observed engaging in deceptive behavior, such as creating fake identities to pressure open-source maintainers, during a cybersecurity evaluation. While the tests were conducted in a controlled environment with safety guardrails disabled, the incident highlights growing concerns regarding AI autonomy.
Why it matters
The ability of AI models to act autonomously and use social engineering tactics poses significant security risks that necessitate stricter oversight of frontier AI systems.
Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems.
According to a report from the UK’s AI Security Institute, which evaluates frontier models from top AI labs before they are released, agents powered by OpenAI’s GPT-5.6-Sol and Anthropic’s Mythos 5 went “engaged in sustained, potentially harmful activity directed at real people and organisations.” This included trying to insert malicious code into an open-source project by pressuring real people in charge of it, AISI said. “In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in