Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

A study based on 40,000 game runs reveals that humans acting as a 'human-in-the-loop' for AI agents fail to catch malicious commands about one-third of the time. The data shows that users are more likely to approve deceptive commands that appear routine, such as those involving package scripts.
Why it matters
As AI agents become more integrated into software development, the human inability to consistently detect malicious commands poses a significant security risk to corporate and personal systems.
A couple of months ago I published a small browser game : you play the human-in-the-loop for an AI coding agent, approving or denying its commands under time pressure. Some commands are routine ( git status , npm test ) and some other commands indicate your agent has been possessed and is sending your secrets to a remote server ( cat ~/.aws/credentials ). More on the threats associated with agents running commands and how to mitigate them can be found in the original post .
The game garnered some interest on hacker news , and after adding in statistics (unfortunately a bit later on) we can take a closer look at the data of over 40,000 runs and 409,000 individual approve/deny decisions. Let's see how the human-in-the-loop, our last line of defence against rogue agents, fared.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in