Key facts
- Strict guardrails on advanced AI models are hindering cybersecurity researchers' ability to find and exploit vulnerabilities.
- U.S. government export controls were temporarily placed on Anthropic's AI models due to concerns about bypassing security measures.
- Cybersecurity experts argue that AI tools are essential for both offensive and defensive security work, and guardrails make them less effective.
- Researchers are increasingly relying on open-source AI models that lack restrictions.
- Some experts believe these guardrails are pushing researchers towards foreign AI systems and could be detrimental to the AI race.
Strict guardrails implemented by AI companies to prevent malicious use of their models are now impeding the work of legitimate cybersecurity researchers, according to industry experts. These restrictions, intended to stop hackers from building and executing cyberattacks, are also hindering network defenders and those who proactively probe systems for weaknesses.
Concerns have been amplified by U.S. government export control restrictions, which were temporarily placed on Anthropic's AI models Mythos and Fable due to reports of guardrails being bypassed. While these specific controls have since been lifted, access to some advanced models remains restricted to vetted users.
Researchers argue that these limitations prevent AI from being used effectively in identifying and exploiting vulnerabilities, a crucial part of defensive cybersecurity. Mark Dowd, a security researcher, expressed discomfort with large companies making arbitrary decisions about security. Chris Anley of NCC Group highlighted that AI tools are dual-use, essential for both offense and defense, and that guardrails disrupt this balance.
As a result, many researchers are turning to open-source AI models that can be run locally without any restrictions. Paolo Stagno of CrowdFense noted that AI companies treat customers like children with their vetting programs. Giuseppe Cali, however, stated that guardrails do not impede his work as he uses AI for reverse engineering and tool building, not for direct vulnerability discovery or weaponization.
Chris Thompson of RemoteThreat pointed out that guardrails can be inconsistent and lead to researchers spending more time negotiating with models than on core security tasks. He warned that responsible researchers are being pushed towards foreign-owned AI systems, potentially harming the global AI race and leaving defenders unprepared for future attacks.