Mythos Vulnerability Firehose Hits a Human Bottleneck

A staggering 90% of vulnerabilities discovered by Anthropic’s Claude Mythos AI model have failed to reach the disclosure stage, and only a tiny fraction of those that did have been fixed. This alarming gap highlights a human bottleneck in validating new flaws and coordinating their remediation, rather than in discovering them.

The analysis, conducted by Patrick Garrity, a security researcher at VulnCheck, examined public data from Anthropic’s Vulnerability Disclosure Ledger, which tracks the progress of Project Glasswing-related findings as they move through the vulnerability disclosure and remediation process. Since its launch in April 2026, Claude Mythos has generated a whopping 26,153 vulnerability findings across numerous software projects. However, only 2,736 of these – just over 10% – have made it into the disclosure ledger.

This discrepancy is not surprising to those familiar with the complexities of coordinated vulnerability disclosure. But it does challenge the narrative presented by frontier model providers, which has largely focused on AI’s ability to dramatically accelerate vulnerability discovery. As Garrity notes, “It seems like they are learning this through trial and error.” The true extent of AI’s limitations in this area is only now becoming clear.

Garrity’s analysis also raises questions about Anthropic’s claims regarding the accuracy of Mythos’s vulnerability findings and its ability to assess their severity. He points out that the 202 findings marked as fixed are notably fewer than the 245 vulnerabilities marked as withdrawn, which warrants closer scrutiny of how Anthropic defines and measures its claimed 91.4% true-positive rate.

Furthermore, Garrity found that Claude Mythos is significantly more aggressive in assessing the severity of vulnerabilities compared to the actual maintainers of the affected software. The AI model assessed 91.5% of the findings as critical or high severity, but project maintainers determined only 61.3% as being in this category. This discrepancy suggests that Anthropic’s team may not have provided clear instructions on how to determine severity, resulting in inflated determinations.

The issues surrounding Claude Mythos are part of a larger problem affecting AI-powered security scanners. Research by Contrast Security has shown significant variability in the results produced by these tools, with different runs against the same codebase producing substantially different findings and multiple scanners agreeing on only a small percentage of vulnerabilities.

This raises important questions about the accuracy and reliability of AI-generated vulnerability findings. As Jeff Williams, founder of Contrast Security, notes, “When my company ran three different AI scanners three times against the same 50,000 line codebase, the scanners generated different results.” This highlights the need for a more nuanced understanding of AI’s limitations in this area.

So what does this mean for users and developers? It’s essential to approach AI-generated vulnerability findings with a critical eye, recognizing that human validation and coordination are crucial steps in determining which flaws warrant disclosure and remediation. As Garrity notes, “The numbers warrant closer scrutiny of how Anthropic defines and measures its claimed 91.4% true-positive rate.” By acknowledging the limitations of AI-powered security scanners, we can better navigate the complex landscape of vulnerability research and ensure that only credible findings make it to the disclosure stage.


Source: Dark Reading — 2026-09-09