Project Glasswing: What Anthropic's AI Cyber Evaluations Mean for Security Validation

Project Glasswing is an Anthropic controlled initiative for using advanced AI models to find and help fix vulnerabilities in critical software before similar capabilities become widely available to attackers. It has drawn attention because Claude Mythos Preview did more than identify suspicious code. It reproduced vulnerabilities, developed exploit components and combined weaknesses into complete attack chains.
One point needs clarification: Project Glasswing is a restricted research and defense program, and it's often mistaken for a single benchmark with one score. The program combines live code reviews and partner deployments. Independent evaluations such as CyberGym and ExploitBench add measured evidence of what the model can do. Together, they show how frontier AI models are reshaping both cyber offense and cyber defense.
For enterprises, the concern is practical. As AI models improve, attackers may need less time and expertise to discover exploitable weaknesses and develop sophisticated attack chains. Security teams must understand those capabilities, then continuously validate whether their prevention, detection and response controls can withstand the techniques these models help scale. Continuous security validation helps organizations move beyond assumptions and verify that security controls perform effectively against evolving threats.
Key Takeaways
- Project Glasswing is a controlled research program. It combines live vulnerability discovery, exploit development and coordinated disclosure, while independent benchmarks measure how frontier AI models perform across the cybersecurity lifecycle.
- AI is accelerating cyber offense. Advanced models like Claude Mythos can identify vulnerabilities, reproduce exploits and build attack chains, reducing the time and expertise required for sophisticated attacks.
- Organizations need more than vulnerability management. As AI increases the volume and complexity of findings, security teams must continuously validate which exposures are actually exploitable, whether existing security controls can stop them and mitigate with updates to security controls, “virtual patches” if patch is not available.
- Continuous security validation closes the gap. AI capability benchmarks show what frontier models can do. Regularly testing prevention, detection and response controls against realistic attack techniques shows whether an organization's defenses can withstand those capabilities in practice and discovers gaps with specific mitigation.
What Is Project Glasswing?
On April 7, 2026, Anthropic launched Project Glasswing in response to a sharp increase in AI cyber capability. The program is built around Claude Mythos Preview, an unreleased frontier model designed for advanced cybersecurity research.
Claude Mythos is the model, and Project Glasswing is the controlled research and defense program around it. Glasswing supplies the governance, access controls, disclosure processes and partner ecosystem needed to use the model's capabilities for defense.
Those capabilities are significant. Mythos Preview found vulnerabilities in major operating systems, browsers and open-source components, including flaws that had survived years of human review and millions of automated test runs. It could also go from finding a bug to building a working exploit with little human direction.
The initial partners were Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks. Anthropic then gave access to more than 40 more organizations that build or maintain critical software infrastructure. Those organizations can use the model to scan and secure their own systems and open-source software. Anthropic also committed to sharing what it learns so the whole industry can benefit.
Anthropic later released Mythos 5, an updated model that, as of October 2026, is available only to vetted partners.
Project Glasswing at a glance
| Area | Summary |
| Creator | Anthropic, working with technology companies, critical infrastructure providers and open-source maintainers |
| Core model | Claude Mythos Preview, later followed by Mythos 5 |
| Primary purpose | Find, validate, disclose and help remediate vulnerabilities in critical software |
| Evaluation approach | Live code analysis, controlled exploit development and coordinated disclosure, supported by results on independent benchmarks such as CyberGym and ExploitBench |
| Main capabilities tested | Vulnerability discovery, exploit reproduction, multi-step reasoning, tool use and attack-chain construction |
| Scope limits | A single public leaderboard, a general measure of cyber risk or a test of one enterprise's defenses all fall outside its scope |
| Why it matters | It shows how frontier models may lower the cost and expertise needed for sophisticated cyber operations |
Anthropic expanded the program in June 2026 to approximately 150 additional organizations across more than 15 countries. The expansion included power, water, healthcare, communications and hardware providers, underscoring the growing importance of AI-assisted cybersecurity research for critical infrastructure.
Why AI Cybersecurity Benchmarks Matter
General AI benchmarks can measure knowledge, coding performance or reasoning. They don't show whether a model can operate through the full lifecycle of a cyber task.
A model may score well on a coding benchmark yet fail to recognize a trust boundary, compile a proof of concept or adapt after an exploit attempt fails. Cybersecurity evaluations must test actions as well as answers. The model needs to inspect unfamiliar code, form a hypothesis, use tools, interpret runtime behavior and revise its plan.
That is why tests such as MMLU or SWE-bench provide only part of the picture. MMLU measures broad academic knowledge. SWE-bench tests whether models can resolve software issues in real repositories. Neither was built to determine whether a model can turn a memory-safety flaw into a reliable exploit chain.

Specialized AI cyber evaluations measure these higher-risk capabilities under controlled conditions. They also help researchers identify thresholds that may require stronger safeguards or restricted access.
The need is growing because AI supports both sides of security operations. Defenders use it for code review, vulnerability triage, detection engineering and incident analysis. Attackers can apply similar capabilities to reconnaissance, phishing, exploit development and post-compromise activity. As agentic AI becomes more capable of planning, using tools and adapting based on results, periodic security assessments can't keep up. Organizations need continuous validation to confirm that prevention, detection and response controls remain effective against rapidly evolving attack techniques.
Standardized evaluations give developers, governments and security leaders a common evidence base. They help separate marketing claims from tested performance and support more responsible release decisions.
How Project Glasswing Works
Project Glasswing combines controlled model access, authorized testing and coordinated vulnerability disclosure. Partners apply Mythos to software they own or are allowed to assess. The model can review source code, generate test cases, compile code and run proof-of-concept attempts inside isolated environments.
Cloudflare described a multi-stage harness it used with Mythos. One group of agents mapped repositories and attack surfaces. Other agents hunted for narrowly scoped vulnerability classes. Independent agents tried to disprove findings. Later stages removed duplicates, traced whether attacker-controlled input could reach a flaw and produced structured reports.
This structure matters. A single prompt can produce inconsistent results. A harness divides work, validates output and preserves evidence.
Independent benchmarks, built outside the program, add controlled measurement of the same capabilities:
- CyberGym tests vulnerability reproduction in real software.
- ExploitBench measures exploit development across 41 patched V8 vulnerabilities and 16 capability levels.
- ExploitGym evaluates exploit creation across vulnerabilities from OSS-Fuzz, V8 and the Linux kernel.
- SCONE-bench focuses on smart contract exploitation.
These evaluations test reasoning, planning and execution. The model must reach vulnerable code, reproduce the bug, create exploit primitives and sometimes achieve unauthorized code execution. That is a much harder task than recalling a known common vulnerabilities and exposures (CVE) description.
Project Glasswing measures model capability in authorized environments. It doesn't prove that a model can compromise any live organization, and it doesn't account for every identity control, network path, cloud permission or detection layer. It also doesn't test whether an organization's controls would stop the resulting techniques.
Enterprise security validation must answer the next question: What happens when comparable techniques meet the organization's actual defenses?
Key Findings from Project Glasswing
Anthropic reported that approximately 50 initial partners used Mythos Preview to find more than 10,000 high- or critical-severity vulnerabilities. Several partners said their discovery rate increased by more than 10 times. Cloudflare reported 2,000 bugs across critical systems, including 400 rated high or critical.
Anthropic also scanned more than 1,000 open-source projects and reported 23,019 potential vulnerabilities. Mythos estimated that 6,202 were high or critical. Independent researchers reviewed 1,752 of those findings. Of the reviewed group, 90.6% were valid true positives and 62.4% were confirmed as high or critical. At the time of the update, Anthropic had reported 530 high- or critical-severity bugs to maintainers, and 75 of them had been patched.

The model performed strongly on formal evaluations. Mythos Preview scored 83.1% on CyberGym, compared with 66.6% for Claude Opus 4.6. In Epoch AI's review of ExploitBench, Mythos Preview achieved arbitrary code execution on 16 to 18 of the 41 V8 vulnerabilities, depending on the test harness.
The case studies show why these scores matter. Mythos found a 27-year-old OpenBSD vulnerability, a 16-year-old FFmpeg flaw and a Linux kernel privilege-escalation chain. Mozilla reported finding and fixing 271 vulnerabilities in Firefox 150 while testing the model, more than 10 times the number it found in Firefox 148 using Claude Opus 4.6.
The limitations remain important. Models can generate noise, miss context or produce patches that create regressions. Cloudflare found that structured validation and independent review were necessary. Anthropic also reported that the main bottleneck shifted from finding vulnerabilities to verifying, disclosing and patching them.
Frontier models have crossed an important threshold, and expert oversight still matters. Human teams need to confirm impact, coordinate disclosure, prioritize remediation and verify fixes. For enterprise security teams, AI may accelerate vulnerability discovery, but organizations still need continuous validation to understand which findings create real exposure in their own environments.
What Project Glasswing Means for Enterprise Security
For security leaders, the main change is the speed, scale and accessibility of existing attack methods, rather than a new category of attack.
AI can help more people analyze code, reproduce vulnerabilities and connect smaller weaknesses into exploitable paths. That may shorten the time between discovery and weaponization. It can also increase the volume of plausible findings that security teams must assess.
Vulnerability management must therefore become more context-aware. As AI produces more findings, organizations need to distinguish theoretical vulnerabilities from those that are actually exploitable within their own environments. Exposure validation provides that context by combining attack simulation with evidence about how security controls perform in practice.
Compensating controls matter more when discovery outpaces patching. Segmentation, least privilege, application isolation and strong identity controls can limit impact before a patch is available with fixes to security controls, known as virtual patches. Continuous exposure validation verifies that these controls perform as intended, so teams don't have to assume they will during a real attack and delivers mitigations for gaps. Cloudflare reached a similar conclusion after testing Mythos, noting that faster patching alone doesn't solve the architectural problem around vulnerable systems.
Detection and response content must also keep pace. AI-assisted operators may vary payloads, combine techniques and retry faster. Static assumptions and annual testing can't provide enough confidence.
Organizations also need governance around their own use of cyber AI. Approved agents should operate within defined permissions, logging and human review. Unapproved AI tools can expose source code, credentials or sensitive data, making AI governance an increasingly important component of enterprise security. For more on this topic, see our guide to combating rogue AI.
From AI Benchmarks to Continuous Exposure Validation
Project Glasswing helps answer one question: Can frontier AI find and exploit difficult software weaknesses?
Enterprises need to answer another question: Can our defenses prevent, detect and respond to the techniques at the same speed and scale that capable models help attackers execute?

That requires continuous exposure validation that leverages agentic AI. Instead of relying only on vulnerability counts or point-in-time assessments, security teams safely reproduce realistic attack activity and observe how controls respond. They can test common security defenses such as endpoint detection and response (EDR), security information and event management (SIEM), web application firewalls (WAF), email and web gateways, identity controls and network defenses.
The process should validate three outcomes: whether a control prevented the action, whether the organization generated useful telemetry for detection and whether responders could contain the activity before it reached a critical asset.
This evidence improves prioritization. A high-severity vulnerability may present limited risk if it is unreachable and protected by effective controls. A lower-severity weakness may require urgent action if it connects to exposed credentials, excessive permissions or an attack path to a critical system.
Cymulate Exposure Validation leverages agentic cyber defense engineering to continuously tests threats, techniques and security controls in production-safe ways, using a daily feed of new threats and MITRE ATT&CK techniques. Combined with Cymulate Detection Studio for ingesting and validating existing SIEM rules and Cymulate CTEM for prioritizing exploitable exposure, the Cymulate Platform helps teams validate prevention and detection across their existing security stack.
We go beyond validation. The cyber defense engineering control plane in the Cymulate Platform understands how each control responded to an attack and builds vendor-specific mitigations, such as detection rules, indicators of compromise (IoCs) and recommended control updates. Cymulate Auto Mitigation can then push those updates to the security stack.
This supports a shift toward agentic cyber defense engineering. Agentic cyber defense engineering is a closed-loop, AI-assisted approach to continuously proving and improving security defenses. Agentic cyber defense engineering uses AI agents, exposure validation and security control integrations to continuously test, tune and improve cyber defenses at machine speed.

Glasswing shows what advanced models can allow attackers to do. Continuous exposure validation that leverages agentic AI shows whether an organization can withstand it and respond at machine speed, as needed.
Conclusion
Project Glasswing marks an important milestone in AI cybersecurity evaluation. It shows that frontier models can discover difficult vulnerabilities, develop exploit components and connect weaknesses into attack chains at a scale that changes long-standing assumptions about vulnerability research.
It also shows the limits of capability benchmarks. A strong model score doesn't reveal whether a specific enterprise is exposed, whether controls will prevent or detect the activity as well as whether responders can contain it.
Security teams need both perspectives. Benchmarks show what advanced AI may enable. Continuous exposure validation that leverages agentic cyber defense engineering proves how their own defenses perform against evolving techniques and implements mitigations, all at machine speed.
As AI-assisted operations become faster and more accessible, confidence must come from tested evidence. Organizations that continuously prove, prioritize and adapt their defenses can find gaps earlier, remediate by real risk and update controls with virtural patches before a capability shift becomes an incident.
See how Cymulate Exposure Validation proves what's exploitable in your environment and builds the mitigations with security control updates to close the gap so your organization can be ready for attackers executing at machine speed.