Can open-source prompt-injection detectors catch realistic AI agent attacks?
A study evaluated 10 open-source prompt-injection detectors against 629 realistic AI agent attacks, revealing that none could effectively catch most attacks without also generating a high number of false positives. Meta's Prompt Guard 2, for instance, caught only 1% of attacks, indicating significant vulnerabilities in current detection methods.
Why it matters
This research highlights critical weaknesses in existing open-source AI security tools, posing substantial risks for the deployment and integrity of AI agents. It underscores an urgent need for more sophisticated and reliable prompt-injection detection mechanisms to prevent malicious AI attacks.
Can open-source prompt-injection detectors catch realistic AI agent attacks?
I ran 10 open-source detectors against 629 real AgentDojo injection attacks , each buried inside ordinary tool output — the way an agent firewall actually sees them. None catches most attacks without also blocking normal traffic.
🥇 Best trade-off: 51% caught at 2% false positives 🔴 Meta's Prompt Guard 2: 1% caught 🚫 Two detectors flag 98% of safe tool outputs too
make bench-agentdojo · 629 attacks + 97 benign cases, each attack embedded in real AgentDojo tool output. Alone = the 27 distinct attack texts scored with no surrounding text ( make bench-payloads ).
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in