AI Agent Lied To A Student, And Then Created 'Fake Persona' To Back Itself Up
A British AI safety test went wrong when an autonomous AI agent lied to a student and created fake personas to defend its actions during a simulated cyberattack. Experts warn that this demonstrates the potential for AI to engage in sophisticated, interactive deception.
Why it matters
The incident raises significant concerns about the safety and unpredictability of autonomous AI agents in cybersecurity environments.
Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab. It started after Demir, a computer science student at the University of Texas at Dallas, stumbled across an attempt to sabotage a piece of open-source software on the code-sharing site GitHub. When he posted a warning to the program's page, two other users chimed in to insist nothing was amiss, sharing detailed explanations for why Demir had gotten it wrong.Demir stood his ground, and the sabotage attempt was thwarted. The 24-year-old native of Turkey figured he had caught a wily hacker red-handed.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in