Microsoft Says New Cybersecurity AI Model Helps MDASH Hit 95.95% at Half the Cost
Microsoft has launched MAI-Cyber-1-Flash, a new cybersecurity-specific AI model integrated into its MDASH vulnerability management system, achieving a 95.95% score on CyberGym at 50% less cost.
Intelligence analysis by Gemini 2.5 Flash

The new AI model, MAI-Cyber-1-Flash, is designed to handle 90% of MDASH's vulnerability identification and remediation tasks, with GPT-5.4 reserved for the remaining complex 10%. Microsoft claims this configuration significantly boosts performance on known-vulnerability tests while substantially reducing operational costs for approved MDASH customers.
Imagine a super-smart robot helper for fixing computer problems. Microsoft built a new, faster, and cheaper brain for this robot that can find most problems quickly, like a junior detective. For the really tricky cases, it still asks a super-expert robot for help. This makes fixing computer security issues much faster and cheaper, like having a whole team of detectives working for less money.
Analysis
MDASH's AI Evolution
Microsoft has introduced MAI-Cyber-1-Flash, its inaugural cybersecurity-specific AI model, as a core component of MDASH, the company's multi-model vulnerability identification and remediation platform. This new model is a sparse mixture-of-experts transformer, boasting 137 billion total parameters with five billion active parameters and a substantial 256,000-token context window. It represents a cybersecurity fine-tune of MAI-Code-1-Flash, itself derived from a MAI-Thinking-1 mid-training checkpoint.
The strategic design behind MAI-Cyber-1-Flash is its role in task routing within MDASH. Microsoft intends for this smaller, more specialized model to manage up to 90% of MDASH's tasks, reserving the more powerful GPT-5.4 for the most challenging 10%. This integration aims to optimize resource allocation and efficiency, though MAI-Cyber-1-Flash is exclusively available within MDASH and not as a standalone public model or API.
Benchmarking and Cost Efficiency
Microsoft reports that MDASH, utilizing MAI-Cyber-1-Flash alongside GPT-5.4, achieved a 95.95% score on CyberGym Level 1. It's crucial to note that CyberGym Level 1 is a known-vulnerability reproduction test, assessing an agent's ability to generate a working proof of concept from a vulnerability description and unpatched code, rather than blind vulnerability discovery or patch correctness. The reported score was not listed on CyberGym's public leaderboard at the time of the article, which still showed Microsoft's earlier 88.4% submission.
Furthermore, Microsoft claims a 50% cost saving compared to its previous best MDASH configuration, which included GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. However, the announcement and model card do not disclose the underlying metrics—such as token use, call volume, latency, or compute allocation—necessary for independent reproduction or normalization of this cost comparison. The model card also reports varying scores on other benchmarks like CVEBench (0.314), CyberSecEval4 (0.553), and a malware-analysis test (0.33), but notably scored zero across ExploitGym's kernel, userspace, and browser categories, which require generating code-execution exploits.
Broader Strategic Implications
This launch marks the first scenario announced for Project Perception, Microsoft's overarching initiative to coordinate defensive security agents. Taesoo Kim, Microsoft's vice president of agentic security, emphasized the distinction that "the model is one input, the system around it is the product," highlighting MDASH as the comprehensive solution. Project Perception is slated for a public preview on August 3, with plans to expand the model's capabilities beyond software vulnerability management to encompass additional security workflows.
While the advancements are promising, Microsoft's model card includes a warning that generated text and code may be inaccurate or incomplete and necessitates review before consequential use. All benchmark testing was conducted in a network-isolated environment, without access to production systems or the public internet, underscoring a cautious approach to deployment. This strategic move by Microsoft signals a significant push towards AI-driven automation in cybersecurity, aiming to enhance defensive capabilities and streamline security operations for its customers.
Key points
- Microsoft launched MAI-Cyber-1-Flash, a new cybersecurity-specific AI model for its MDASH vulnerability management system.
- MDASH, with the new model, scored 95.95% on CyberGym Level 1, a known-vulnerability reproduction test.
- The new configuration is claimed to reduce costs by 50% compared to previous MDASH model mixes.
- MAI-Cyber-1-Flash handles 90% of tasks, with GPT-5.4 reserved for the hardest 10%, and is not available standalone.
- This is the first scenario for Project Perception, Microsoft's broader system for coordinating defensive security agents, with public preview scheduled for August 3.
This new AI model could significantly enhance the efficiency and reduce the cost of vulnerability management for organizations using MDASH, potentially leading to faster identification and remediation of security flaws. Its integration into Project Perception suggests a broader future where AI agents coordinate to provide comprehensive defensive security, making systems more resilient against attacks.
The reported benchmark results have significant caveats, including the test's nature (known vulnerabilities) and the lack of independent verification for cost savings or performance comparisons. Over-reliance on AI-generated fixes without human review, as warned by Microsoft, could introduce new vulnerabilities or fail to address complex threats effectively, potentially creating a false sense of security.



