MDASH tops CyberGym with 95.95% score
MDASH, using Microsoft’s MAI-Cyber-1-Flash, scored 95.95% on CyberGym, outpacing Mythos 5 and GPT-5.6 Sol. MDASH is in private preview through Microsoft Defender.
Microsoft reported that MDASH, paired with a new cybersecurity model called MAI-Cyber-1-Flash, reproduced 95.95% of known vulnerabilities on the CyberGym benchmark. CyberGym asks AI agents to recreate 1,507 documented vulnerabilities drawn from 188 open-source projects and scores success by the percentage of vulnerabilities reproduced in a controlled environment.
The company described MAI-Cyber-1-Flash as handling up to 90% of the workload in the combined system. MDASH routes the most difficult 10% of cases to GPT-5.4 to confirm findings and build proof-of-concept exploits, a configuration Microsoft says reduces token costs by using an efficient model for the bulk of work.
Microsoft compared the result to recent public CyberGym performances. The company reported MDASH’s 95.95% score ahead of GPT-5.5 Cyber at 85.6%, Anthropic’s Mythos 5 at 83.8%, OpenAI’s GPT-5.6 Sol at 83.6% and Gemini 3.5 Flash Cyber at 83.2%. The firm noted the score is self-reported and had not appeared on CyberGym’s public leaderboard at the time of the announcement.
MDASH is a harness that coordinates more than 100 specialized agents to audit code, challenge suspected findings, remove duplicates and construct working demonstrations that a vulnerability can be triggered. Microsoft said the agent-driven workflow ranks findings and produces artifacts developers can review.
The company said pairing a model built specifically for cybersecurity with an agent-driven system improves efficiency and lowers cost. Satya Nadella wrote on social media: “When combined with MDASH, (MAI-Cyber-1-Flash) delivers world-class performance at 50 percent of the cost of leading models.”
MDASH is available in a private preview through Microsoft Security Exposure Management in the Defender portal. Customers in the preview can scan Git repositories, view findings ranked from unlikely to proven, and generate proposed code fixes via the Defender CLI. The preview limits repositories to roughly 256MB and permits one concurrent scan per tenant.
Microsoft described MAI-Cyber-1-Flash as its first model built specifically for cybersecurity tasks and said it used the company’s security telemetry and historical data in training and validation.
External researchers and security teams have replicated similar vulnerability-hunting techniques with public models at lower cost. Researcher Dawid Moczadło observed that “the moat is moving from model access to validation,” referring to the increasing importance of systems that can prove findings without producing large numbers of false positives.
Microsoft said it plans to extend the multi-agent approach beyond code scanning into broader threat detection and remediation under Project Perception. The company added that continued development and validation will take place during the private preview before any wider release.
The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.








