Cybersecurity researchers have been trying to create AI-powered defense systems, but it’s a tough challenge. The problem is that most AI models are much better at attacking systems than defending them. To level the playing field, researchers have started using “red team” agents to help teach their “blue team” counterparts.
Red and blue teams refer to two types of agents: the red team simulates attackers trying to breach a system, while the blue team tries to defend against those attacks. But it turns out that creating effective defense systems is much harder than attacking them. The reason is that defensive tasks require nuanced skills that are difficult to measure or replicate.
To make matters worse, most benchmark tests for AI-powered security agents focus on offense rather than defense. These tests can give a good idea of how well an agent performs in a specific task, but they don’t provide the data needed to train the blue team agents effectively. This creates a bottleneck: generating enough security data to train the blue team agents requires significant time and effort.
To address this issue, researchers at Dreadnode, an AI offensive security startup, have developed two open-source tools that can help evaluate the effectiveness of agentic defenders. The first tool is called DreadGOAD, which simulates Active Directory environments and mimics the messy deployments common in large organizations. The second tool, Ares, is a red team-blue team system designed to test and study offensive and defensive effectiveness.
The researchers used these tools to run automated tests between red and blue team agents. The results were revealing: while the red team agents excelled at simulating attacks, the blue team agents struggled to defend against them. In fact, the researchers initially found that their own blue team agent was “really bad” at defending against simulated attacks.
To improve the performance of their blue team agent, the researchers used a clever trick: they had the red team agents generate data through attack simulations and then applied that data to the blue team agents. The results were encouraging: the blue team agent showed significant improvement in its ability to defend against simulated attacks.
The implications of this research are far-reaching. By leveraging the strengths of both red and blue teams, AI-powered defense systems can become more effective at protecting networks from cyber threats. Moreover, by addressing the uneven playing field between offense and defense, researchers can create more realistic benchmark tests for AI-powered security agents.
For organizations looking to implement AI-powered security solutions, this research provides a valuable takeaway: it’s not enough to simply deploy an agent without proper training or testing. Instead, companies should invest in creating robust benchmark tests that account for both offensive and defensive capabilities. By doing so, they can develop more effective defense systems that stay ahead of evolving cyber threats.
Ultimately, the development of AI-powered security solutions is a complex challenge that requires collaboration between researchers, organizations, and industries. But by sharing knowledge, data, and expertise, we can create better defense systems that protect networks from cyber threats – and give us all a fighting chance in the ongoing battle against cybercrime.
Source: Dark Reading — 2026-07-29