FLAWED's Flaws and What This Means for Industry Research
Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization.
On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs.
I saw FLAWED when it was released and discussed it with other researchers; we classified it as slop and moved on. What I had not realized at the time was how far 1Password’s distribution had carried it: into news coverage and defender roadmaps. Watching this work obscure more rigorous research from less-resourced groups compelled me to post on Twitter, and the responses to that compelled me to write this blog post.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in