Key facts
- A new security technique uses 'banned topics' to halt AI cyberattacks.
- The method, called a 'context bomb,' leverages AI's built-in safety restrictions.
- Tests showed a significant reduction in successful AI attacks, from 57% to 5%.
- This approach aims to work alongside existing AI detection systems like 'canaries'.
- The technique was tailored to different AI models, using topics like biological weapons for Western models.
A novel cybersecurity strategy involves using an AI's own safety restrictions against it to thwart potential cyberattacks. Researchers at Tracebit, a London-based security firm, have demonstrated that by introducing specific 'banned topics' or 'context bombs' into an AI's exploration of a system, its built-in safeguards can be triggered, halting malicious activity.
This technique builds upon previous research where AI models were given a simulated company network to infiltrate. In those earlier tests, decoys known as 'canaries' alerted defenders to AI intrusion attempts with a median lead of eight minutes before full admin access was achieved. While effective, this warning window is narrow given the speed of AI.
The 'context bomb' method aims to buy more time by directly stopping the AI attack before it even reaches the point where a canary alert would be triggered. These context bombs are short text fragments designed to push an AI attacker towards subjects it is programmed to avoid, such as sensitive political topics or the creation of dangerous weapons, depending on the AI model's origin and restrictions.
In trials involving five leading AI models, including Opus 4.8 and Gemini 3.1 Pro, the success rate of AI reaching admin access dropped dramatically from 57% to 5% when context bombs were employed. Full system compromise fell from 36% to 1%. Even the strongest attacker, Opus 4.8, failed every time when a context bomb was present.
Tracebit acknowledges that this method does not entirely solve the broader issue of prompt injection, which exploits AI's tendency to confuse instructions with data. However, they propose that context bombs and canary alerts can work in tandem to provide a more robust defense against AI-driven cyber threats.
