Microsoft and Wiz mind-meld agents catch more than 90% of bugs

Microsoft and Wiz have developed AI-driven bug-hunting agents that demonstrate high success rates in identifying software vulnerabilities. By utilizing specialized models for specific security tasks, these systems have outperformed general-purpose AI models in automated code analysis.
Why it matters
The success of these agentic systems marks a significant advancement in cybersecurity, potentially automating the detection and remediation of zero-day exploits.
Secret to their success: Using the right model for the right security job
Two agentic bug-hunting systems from Microsoft and Google-owned Wiz show that when it comes to finding and remediating software vulnerabilities, at least two models’ minds work better than one - and Wiz tells us it’s adding a third.
Wiz on Monday said Project Atlas, its bug-hunting AI agent, bested Anthropic’s Mythos Preview and OpenAI’s GPT-5.5 Cyber with its vulnerability-analysis skills, achieving a 90.9 percent success rate on CyberGym, and uncovering more than 200 zero-day security holes in widely used open-source code.
Meanwhile, Microsoft boasted its MDASH bug-hunting harness scored a 95.95 percent success rate on CyberGym, also beating Mythos, Gemini and GPT on the same benchmark for evaluating how well AI systems find real vulnerabilities in the code.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in