AI Safety Crisis: New Approach Peeks Inside “Black Box” of Large Language Models
A growing concern has emerged in the cybersecurity community about the safety of large language models (LLMs), which can be exploited by malicious actors to produce unwanted content. To address this issue, a team of researchers from Ben-Gurion University of The Negev is proposing a novel approach that focuses on identifying specific cognitive elements within LLMs that may indicate when they take an unwanted action.
The current method of analyzing tokenized inputs and outputs can only go so far in detecting malicious activity. Attackers have found ways to evade these defenses by using different languages for prompts or exploiting vulnerabilities in the system. The researchers argue that treating AI systems as a “black box” is not sufficient, and instead propose focusing on what happens inside the model.
The team’s approach, called Governance via Activation-based Verification and Extensible Logic (GAVEL), aims to create an open system of identified cognitive elements and rules that detect specific types of safety events. This method involves instrumenting the model with digital probes that can record activation patterns across various neuron layers. By analyzing these patterns and mapping them to different categories, researchers can identify potential security threats more effectively.
Unlike previous efforts, which labeled broad ranges of activity as “cybercrime” or “misinformation,” GAVEL uses a more granular scheme of cognitive elements (CEs) that can be combined in logical statements. For example, combining CEs such as “create content,” “provide/give,” and “personal information” results in a rule for detecting potential phishing attacks.
The researchers believe that this approach will allow practitioners to focus on high-level aspects of AI behavior, rather than getting bogged down in the low-level math and processes involved in neural networks. By separating these two levels, the team hopes to create a more accessible and user-friendly system for identifying and mitigating security threats in AI.
The importance of addressing AI safety cannot be overstated. LLMs are increasingly being used in various applications, from chatbots to content generation tools. However, their potential for misuse is significant, and without proper safeguards, these systems can be exploited by malicious actors.
To mitigate this risk, organizations should consider adopting the GAVEL approach or similar methods that focus on activation analysis within LLMs. By instrumenting models and analyzing activation patterns, security teams can gain a better understanding of AI behavior and identify potential threats more effectively. This will not only improve the safety of LLMs but also provide a more comprehensive defense against malicious activity in these systems.
Ultimately, the GAVEL approach offers a promising solution to the challenges posed by large language models. By creating an open system of identified cognitive elements and rules, researchers can develop a more nuanced understanding of AI behavior and identify potential security threats before they occur. As AI continues to play an increasingly important role in our lives, addressing its safety concerns is crucial for ensuring the integrity of these systems and protecting against malicious activity.
Source: Dark Reading — 2026-07-28