‘Yellow Teams’ Are Defining the Future of AI Security

As AI models like Claude Mythos and GPT-5.5 gain widespread attention for their potential in cybersecurity, a new breed of engineers is emerging to shape the future of artificial intelligence security. Dubbed “yellow teams,” these innovators are building both defenses against advanced AI attacks and frameworks that attackers will utilize to carry out those assaults.

Yellow teams are not just about developing tools for red team penetration testers or blue team defenders; they’re a crucial part of the cybersecurity ecosystem, working behind the scenes to define what security will look like in years to come. According to Sam Curry, chief information security officer (CISO) at Zscaler and partner of Anthropic’s Project Glasswing initiative, yellow teams are essentially build teams that say: “Hey, Red, what tools could you use to do better attacks?” or turn to Blue and ask: “What would you like on defense that isn’t supplied to you by your vendor community?”

One such example is Netskope CISO James Robinson’s experience with Project Glasswing. When he first got access to Claude Mythos, he was surprised that it wasn’t as easy as point-and-click. The model hedged its findings and reported vulnerabilities that lacked context obvious to a human pen tester. However, the company had already participated in OpenAI’s Daybreak program and organized a small yellow team, Project Red Horizon, which built a “harness” for GPT-5.5. This software cocoon around the AI model defined what it could do, what permissions it had, and what policies and guardrails it must follow.

Building an effective AI harness is no trivial task, as Robinson notes. Even after wrapping GPT-5.5 in a harness, they initially had a ton of findings – mostly due to incorrect context setup, resulting in false positives that needed to be worked through. Effectively focusing the AI’s powers can be complex, which is why companies like Cisco and Microsoft are open-sourcing their own AI harnesses. For instance, Cisco’s Foundry Security Spec restricts an AI model with a detailed “constitution,” employs AI agents for 13 different roles, and has around 130 sub-requirements.

The emergence of yellow teams highlights the growing recognition that cybersecurity is not just about detection and response but also about proactive defense. As AI models become increasingly powerful, it’s essential to have a robust framework in place to harness their potential while minimizing risks. By understanding how AI attacks will unfold and developing corresponding defenses, we can better protect ourselves against the threats of tomorrow.

In practical terms, this means that organizations should prioritize building their own yellow teams or partnering with companies already working on AI security initiatives. This requires investing time and resources into training engineers to develop effective AI harnesses and frameworks for mobilizing the greatest powers of these models while mitigating their greatest threats. As we continue to navigate the ever-evolving landscape of AI-driven cybersecurity, one thing is clear: the future of security will be defined by the innovators building the tools that both attackers and defenders rely on.


Source: Dark Reading — 2026-07-13