HomeEverythingEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
Story archiveAll categories
← All Stories

AI hackers thwarted by 'banned topics' in security tests

Created at 21 Jul · 7:11 AM1 source↑ Market-relevant
IN SHORT

Researchers have developed a method to stop AI hacking agents by triggering their built-in safety protocols, using 'banned topics' that the AI is programmed to avoid. This technique, called a 'context bomb,' significantly reduced successful AI cyberattacks in tests.

✉Newsletter

PiQ Daily

Pick your topics. Get only what matters, on your cadence.

Key Numbers

95.9%canary alert success rate before AI reached admin access
8 minutesmedian lead time from canary alert
14 minutesaverage time for AI to break in fully
6 minutesremaining time for security teams to react
5leading AI models tested
152attack attempts with context bombs
57%attack success rate without context bombs
5%attack success rate with context bombs
36%full compromise rate without context bombs
1%full compromise rate with context bombs
91%any part of attack completion without context bombs
15%any part of attack completion with context bombs

Who's Involved

Tracebit
London-based security firm that developed the 'context bomb' technique
Opus 4.8
AI model that showed significant improvement with context bombs
Gemini 3.1 Pro
AI model tested with context bomb technique
GLM 5.2
AI model tested with context bomb technique
DeepSeek 4 Pro
AI model tested with context bomb technique
Kimi K2.6
AI model tested with context bomb technique
AI hackers thwarted by 'banned topics' in security tests

↳ Why This Matters

This development offers a potential new layer of defense against sophisticated AI-powered cyberattacks by leveraging the inherent safety mechanisms within AI models themselves, potentially providing crucial extra time for human intervention.

Key facts

  • A new security technique uses 'banned topics' to halt AI cyberattacks.
  • The method, called a 'context bomb,' leverages AI's built-in safety restrictions.
  • Tests showed a significant reduction in successful AI attacks, from 57% to 5%.
  • This approach aims to work alongside existing AI detection systems like 'canaries'.
  • The technique was tailored to different AI models, using topics like biological weapons for Western models.

A novel cybersecurity strategy involves using an AI's own safety restrictions against it to thwart potential cyberattacks. Researchers at Tracebit, a London-based security firm, have demonstrated that by introducing specific 'banned topics' or 'context bombs' into an AI's exploration of a system, its built-in safeguards can be triggered, halting malicious activity.

This technique builds upon previous research where AI models were given a simulated company network to infiltrate. In those earlier tests, decoys known as 'canaries' alerted defenders to AI intrusion attempts with a median lead of eight minutes before full admin access was achieved. While effective, this warning window is narrow given the speed of AI.

The 'context bomb' method aims to buy more time by directly stopping the AI attack before it even reaches the point where a canary alert would be triggered. These context bombs are short text fragments designed to push an AI attacker towards subjects it is programmed to avoid, such as sensitive political topics or the creation of dangerous weapons, depending on the AI model's origin and restrictions.

In trials involving five leading AI models, including Opus 4.8 and Gemini 3.1 Pro, the success rate of AI reaching admin access dropped dramatically from 57% to 5% when context bombs were employed. Full system compromise fell from 36% to 1%. Even the strongest attacker, Opus 4.8, failed every time when a context bomb was present.

Tracebit acknowledges that this method does not entirely solve the broader issue of prompt injection, which exploits AI's tendency to confuse instructions with data. However, they propose that context bombs and canary alerts can work in tandem to provide a more robust defense against AI-driven cyber threats.

Frequently asked questions

A 'context bomb' is a piece of text designed to trigger an AI's built-in safety protocols or 'banned topics,' causing it to stop its operation, such as a cyberattack.

In tests, the success rate of AI reaching admin access dropped from 57% to 5%, and full system compromise fell from 36% to 1% when context bombs were used.

No, Tracebit states it does not fix prompt injection entirely but works alongside other detection methods like 'canaries' to improve defense.

Prompt injection is a technique used by attackers to trick AI, whereas a context bomb uses the AI's own safety rules against it to stop an attack.

What Happens Next

01Further research may explore tailoring context bombs to a wider range of AI models.
02Integration of context bomb techniques with existing AI security monitoring systems is likely.

Get the newsletter.

Pick the topics you actually care about. We'll email when there's news worth your time, on the cadence you choose. Cancel any time from your account.

Cadence

How It Developed

Tracebit tested AI models on a simulated company network with vulnerabilities and decoys.
AI models reached admin access in 57% of attacks without intervention.
Tracebit introduced 'context bombs' designed to trigger AI safety rules.
AI attacks dropped to 5% success rate with context bombs.
The technique significantly reduced full system compromise.
Context bombs worked in conjunction with canary alerts for early warnings.

Sources

T1
Could 'banned topics' be the secret to stopping AI hackers?Euronews

Related Stories

Chinese police used AI to infiltrate and monitor dark web
21 Jul · 6:06 AM
Hackers Exploiting Unpatched WordPress Vulnerabilities, Threatening Millions of Websites
20 Jul · 3:56 PM
AI chatbots give inaccurate, unreliable voting advice, study finds
21 Jul · 5:06 AM
AI-generated influencers and imagery are reshaping TikTok Shop, potentially impacting creators
20 Jul · 8:41 PM
Model Context Protocol update aims to simplify AI integration
20 Jul · 9:06 PM