Mythos Vulnerability Firehose Hits a Human Bottleneck

Cybersecurity’s Unholy Trinity: Mythos Vulnerability Firehose Hits a Human Bottleneck

A new analysis of public data from Anthropic’s Project Glasswing has revealed a disturbing trend in the world of vulnerability research. Despite generating an astonishing 26,153 potential security flaws using its Claude frontier model, only a tiny fraction – less than 10% – have made it to the disclosure stage, where they can be addressed by software maintainers.

The numbers are staggering: out of nearly 26,000 potential vulnerabilities, a paltry 2,736 have reached the disclosure ledger. Of those, a mere 202 have been patched, while 245 were withdrawn from consideration altogether. Another 191 remain in the pre-disclosure stage, waiting to be reported to their maintainers. The rest – nearly 90% of Claude Mythos-generated findings – are still languishing in limbo, awaiting human validation and coordination.

This bottleneck is not just a minor hiccup; it has significant implications for the cybersecurity industry as a whole. As AI-powered tools like Claude continue to churn out potential vulnerabilities at an alarming rate, the question arises: how many of these are actually worth fixing? According to Patrick Garrity, a security researcher at VulnCheck, human validation and coordination have become the new chokepoints in determining which AI-generated findings warrant disclosure and remediation.

Garrity’s analysis also raised questions about Anthropic’s claims regarding the accuracy of Mythos’s vulnerability findings. Specifically, he noted that the 202 findings marked as fixed are significantly fewer than the 245 vulnerabilities marked as withdrawn. This discrepancy warrants closer scrutiny of how Anthropic defines and measures its claimed 91.4% true-positive rate.

Furthermore, Garrity found that Anthropic’s AI is overly aggressive in assessing the severity of vulnerabilities. While Claude assessed a whopping 91.5% of findings as critical or high severity, project maintainers themselves determined only 61.3% to be in this category. This disparity highlights the need for more nuanced and accurate assessments of vulnerability severity.

The implications of these findings go beyond just Anthropic’s Project Glasswing. Research by Contrast Security has shown significant variability in the results produced by AI-powered security scanners, with different runs producing vastly different findings. This raises important questions about the reliability and accuracy of AI-generated vulnerability reports.

As the cybersecurity industry continues to grapple with the challenges posed by AI-powered tools, it is essential that we take a step back and assess the validity of these claims. How many of these vulnerabilities are truly worth fixing? Are AI-generated assessments of severity accurate? These questions demand answers, and it’s time for the industry to have an honest conversation about the limitations and potential pitfalls of relying on AI-powered security tools.

For users and developers, this means being cautious when dealing with AI-generated vulnerability reports. Don’t be swayed by the promise of speed and efficiency; instead, take a closer look at the underlying data and methodology used to generate these findings. Verify the accuracy of severity assessments and scrutinize the definitions and metrics used to measure true positives. By doing so, you can ensure that your organization is not falling prey to a vulnerability firehose – and that the human bottleneck becomes an opportunity for improvement rather than a hindrance to progress.


Source: Dark Reading — 2026-09-09