Microsoft Says MDASH Beats Claude Mythos and GPT

Microsoft has introduced a new cybersecurity AI model, MAI-Cyber-1-Flash, integrated into its MDASH vulnerability-hunting system. The company claims this setup outperforms competitors like Anthropic and OpenAI while significantly reducing operational costs.
Why it matters
This development highlights the ongoing arms race in AI efficiency and specialized cybersecurity tools, potentially lowering the barrier for enterprise-grade threat detection.
Microsoft has released its first dedicated cybersecurity model named MAI-Cyber-1-Flash and plugged it into MDASH, a vulnerability-hunting system that it says beats Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol while costing 50% less than Microsoft’s current best MDASH configuration, according to the company.
The combined setup scored 95.95% on CyberGym, according to Microsoft. CyberGym is a benchmark that asks AI agents to reproduce 1,507 known vulnerabilities across 188 open-source projects, then scores them by the percentage successfully reproduced in a controlled environment.
That put MDASH ahead of GPT-5.5 Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 83.2%. The result is self-reported by Microsoft and had not appeared on CyberGym’s public leaderboard at publication time, though the benchmark uses a public test set and a defined success metric.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in