‘Yellow Teams’ Are Defining the Future of AI Security

The Future of AI Security is Being Shaped by “Yellow Teams” Behind the Scenes

In a bid to stay ahead of the curve, a growing number of organizations are quietly building defense and attack tools using artificial intelligence (AI) to test their cybersecurity. These secretive engineering teams, dubbed “yellow teams,” are not only developing defenses against future AI-powered threats but also crafting the very frameworks that attackers will use to carry out those attacks.

At the heart of this phenomenon are initiatives like Project Glasswing, launched by Anthropic in April, and OpenAI’s Daybreak program, which invite organizations to play with cutting-edge AI models like Claude Mythos and GPT 5.5. These models have been used by red teams (penetration testers) to exploit their own companies’ systems, while rival blue teams try to detect and defend against those exploits.

As a result, yellow teams are emerging as the unsung heroes of the cybersecurity landscape. They’re building both mitigations against AI’s greatest threats and frameworks for mobilizing its greatest powers. “Yellow team is a build team,” explains Sam Curry, chief information security officer (CISO) at Zscaler, a Glasswing partner. “They’ll say: ‘Hey, Red, what tools could you use to do better attacks?’ As if they were the engineering department behind a major attacker.”

One of the key challenges facing organizations is effectively harnessing AI’s powers without giving attackers an open invitation. When Netskope CISO James Robinson first got access to Mythos, he was surprised by its limitations. The model reported vulnerabilities that lacked context obvious to human pen testers, highlighting the need for a more nuanced approach.

To overcome these hurdles, companies are building “harnesses” around AI models like GPT-5.5 and Claude Mythos. A harness is essentially a software cocoon that defines what an AI can do, what permissions it has, and what policies and guardrails it must follow. By doing so, organizations can focus the AI’s powers without unleashing its full potential on their networks.

Building an effective AI harness is no trivial task. Netskope’s Robinson remembers that even after wrapping GPT-5.5 in a harness, they had to work through a ton of false positives. But with the right approach, these harnesses can become a crucial step toward effective AI-driven vulnerability hunting.

Companies like Cisco and Microsoft have already open-sourced their own AI harnesses, while Cloudflare has developed its own eight-step procedure for agents working within the framework. These efforts demonstrate that harnessing AI’s powers is not just a technical challenge but also an organizational one, requiring collaboration between different teams and stakeholders.

As the cybersecurity landscape continues to evolve at breakneck speed, the work of yellow teams will be crucial in defining what security looks like in the years to come. By building both defenses against AI-powered threats and frameworks for mobilizing its powers, these teams are ensuring that organizations stay ahead of the curve – but also that they don’t inadvertently create new vulnerabilities along the way.

For those looking to follow suit, the key takeaway is simple: effective AI-driven vulnerability hunting requires more than just throwing a model at the problem. It demands a nuanced understanding of the technology’s limitations and capabilities, as well as a willingness to collaborate across teams and stakeholders. By embracing this approach, organizations can unlock the full potential of AI while minimizing its risks – and stay ahead of the threats that are shaping the future of cybersecurity.


Source: Dark Reading — 2026-07-13