HomeAll NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
Story archiveAll categories
← All Stories

AI guardrails hinder cybersecurity researchers, experts say

Created at 24 Jul · 1:11 AM1 source↑ Market-relevant
IN SHORT

Strict guardrails on advanced AI models, designed to prevent misuse by malicious actors, are now impeding the work of legitimate cybersecurity researchers and network defenders, according to industry experts. These restrictions can prevent AI from being used effectively in identifying and exploiting vulnerabilities.

✉Newsletter

PiQ Daily

Pick your topics. Get only what matters, on your cadence.

Key Numbers

5Fable version number

Who's Involved

Mark Dowd
Security researcher specializing in zero-day vulnerabilities
Chris Anley
Chief Scientist at NCC Group
Paolo Stagno
Chief Technology Officer at CrowdFense
Giuseppe Cali
Security researcher finding zero-days and developing exploits
Chris Thompson
CEO of RemoteThreat and founder of Offensive AI Con
Anthropic
AI company with restricted models and a Cyber Verification Program
OpenAI
AI company offering Trusted Access for Cyber program

↳ Why This Matters

The effectiveness of cybersecurity defenses may be compromised if legitimate researchers are unable to utilize advanced AI tools to discover and address vulnerabilities before malicious actors exploit them.

Key facts

  • Strict guardrails on advanced AI models are hindering cybersecurity researchers' ability to find and exploit vulnerabilities.
  • U.S. government export controls were temporarily placed on Anthropic's AI models due to concerns about bypassing security measures.
  • Cybersecurity experts argue that AI tools are essential for both offensive and defensive security work, and guardrails make them less effective.
  • Researchers are increasingly relying on open-source AI models that lack restrictions.
  • Some experts believe these guardrails are pushing researchers towards foreign AI systems and could be detrimental to the AI race.

Strict guardrails implemented by AI companies to prevent malicious use of their models are now impeding the work of legitimate cybersecurity researchers, according to industry experts. These restrictions, intended to stop hackers from building and executing cyberattacks, are also hindering network defenders and those who proactively probe systems for weaknesses.

Concerns have been amplified by U.S. government export control restrictions, which were temporarily placed on Anthropic's AI models Mythos and Fable due to reports of guardrails being bypassed. While these specific controls have since been lifted, access to some advanced models remains restricted to vetted users.

Researchers argue that these limitations prevent AI from being used effectively in identifying and exploiting vulnerabilities, a crucial part of defensive cybersecurity. Mark Dowd, a security researcher, expressed discomfort with large companies making arbitrary decisions about security. Chris Anley of NCC Group highlighted that AI tools are dual-use, essential for both offense and defense, and that guardrails disrupt this balance.

As a result, many researchers are turning to open-source AI models that can be run locally without any restrictions. Paolo Stagno of CrowdFense noted that AI companies treat customers like children with their vetting programs. Giuseppe Cali, however, stated that guardrails do not impede his work as he uses AI for reverse engineering and tool building, not for direct vulnerability discovery or weaponization.

Chris Thompson of RemoteThreat pointed out that guardrails can be inconsistent and lead to researchers spending more time negotiating with models than on core security tasks. He warned that responsible researchers are being pushed towards foreign-owned AI systems, potentially harming the global AI race and leaving defenders unprepared for future attacks.

Frequently asked questions

AI guardrails are restrictions and safety measures implemented in AI models to prevent their use for malicious purposes, such as building cyberattacks or exploiting software vulnerabilities.

Researchers argue that these guardrails can prevent them from effectively using AI tools to discover and test software vulnerabilities, which is essential for developing robust defenses.

'Zero days' are previously unknown software flaws and the exploits that take advantage of them, which are highly valued by governments for intelligence operations.

Many researchers are turning to open-source AI models that can be run locally and do not have usage restrictions or guardrails.

What Happens Next

01Anthropic and OpenAI may adjust their vetting processes and guardrail policies.
02Further government reviews of AI export controls and security protocols are possible.
03The cybersecurity community will continue to evaluate the impact of AI guardrails on their work.

Get the newsletter.

Pick the topics you actually care about. We'll email when there's news worth your time, on the cadence you choose. Cancel any time from your account.

Cadence
CME Headlines
  • Is AI Making Inflation Better or Worse?
    22 Jul · 3:26 PM

How It Developed

AI companies implemented guardrails to prevent malicious use of models.
U.S. government imposed export controls on Anthropic's AI models due to concerns about bypassing security measures.
These controls were later lifted, but access to some models remains restricted.
Cybersecurity researchers criticize these guardrails for hindering their work in finding vulnerabilities.
Experts note that AI tools are dual-use, serving both offensive and defensive cybersecurity purposes.
Researchers are increasingly turning to open-source AI models without restrictions.
Some researchers argue that guardrails lead to time spent negotiating with models rather than focusing on security tasks.
Concerns are raised that responsible researchers are being pushed towards foreign-owned AI systems.
Sponsored

London Quick Take - 22 July - UK inflation softens, oil rises and chips rally ahead of Alphabet, Tesla earnings

SAXO

Sources

T1
How AI guardrails are impeding the work of offensive cybersecurity researchersTechCrunch

Related Stories

OpenAI reports AI models breached human control, sparking safety concerns
23 Jul · 2:06 PM
AI companies push for business integration amid model control concerns
23 Jul · 12:00 PM
US Officials Accuse Chinese AI Firm of Stealing Tech; Experts Skeptical
23 Jul · 11:16 AM
OpenAI AI models breached Hugging Face systems during security test
23 Jul · 5:21 AM
Anthropic enhances Claude's voice mode with advanced models and app integration
23 Jul · 7:31 PM