Anthropic’s ambitious push to make its advanced AI models more accessible to cybersecurity defenders has just taken a significant leap forward. The company is expanding access to its Mythos 5 model, which offers powerful capabilities for detecting and mitigating cyber threats, while also unveiling a $35 million open source fund to support the security of vulnerable projects.
At the heart of Anthropic’s strategy is the concept of “guardrails” – limitations on direct interaction with advanced AI models that can be misused by malicious actors. By providing access to specific defensive outputs, such as patches or security alerts, defenders can take advantage of Mythos-class capabilities without exposing themselves to risk. This approach has been tested in Anthropic’s Project Glasswing initiative, launched earlier this year, which gave a select group of organizations early access to the Claude Mythos Preview and its successor, Mythos 5.
Now, Anthropic is taking this concept further by integrating Mythos 5 into security operations, incident response, and detection tools used by teams protecting critical infrastructure. End users won’t interact with Mythos directly; instead, they’ll work through purpose-built interfaces that run the model in the background and return a defined output, such as a list of suggested patches, with abuse-prevention checks to keep the model within its scope.
One notable development is the expansion of Claude Security, which now runs its codebase scans on Mythos 5. Scans will surface each finding with a Common Weakness Enumeration (CWE) category, confidence and severity ratings, and a suggested fix – although any fix still requires human approval before deployment. This means defenders can tap into the capabilities of Mythos 5 without exposing themselves to potential misuse.
In addition to these technical advancements, Anthropic is launching the Defender Advantage Fund ($0xDAF), which will provide $35 million in Claude credits to organizations that help open source maintainers secure their projects. This initiative follows previous support from Project Glasswing, including direct donations and coordinated efforts with other security teams. Grants will be awarded for patching live vulnerabilities, building reusable scanning and patching processes, and pursuing security approaches designed to resist entire classes of attack.
Finally, Anthropic is expanding its Cyber Verification Program, which already offers reduced safeguards on Claude Opus and Sonnet models for authorized security work. The program will now cover broader dual-use capabilities on these models, including vulnerability triaging and validation, with Mythos-class access to follow in the coming weeks.
For security teams looking to take advantage of these developments, Anthropic is encouraging them to apply to the Cyber Verification Program now for reduced safeguards on Opus and Sonnet. Further details on the broader rollout are expected in the coming weeks.
As AI models become increasingly powerful and accessible, the need for robust guardrails and responsible access controls has never been more pressing. By expanding access to Mythos 5 while maintaining its guardrails, Anthropic is taking a significant step towards making advanced AI capabilities available to defenders without exposing them to unnecessary risk.
Source: SecurityWeek — 2026-08-24