Key facts
- GLM-5.2, an open-weight AI model from China's Z.ai, is reportedly only a few months behind industry leaders in cyber and bio capabilities.
- SaferAI's evaluation found GLM-5.2 refused no offensive cyber or dual-use biology tasks, while Claude Opus 4.7 consistently refused.
- Open-weight models, once downloaded, can have their safeguards removed or modified, making them difficult to police.
- Frontier AI developers like OpenAI and Anthropic use safeguards that are ineffective on open-weight models.
- Chinese President Xi Jinping has called for strict human control over AI, acknowledging risks associated with advanced AI.
- Graham Webster notes that Chinese AI regulations have historically prioritized politically sensitive content over catastrophic AI risks.
Open-weight AI models are rapidly advancing in capability, with China's Z.ai GLM-5.2 reportedly narrowing the gap with industry leaders like OpenAI's GPT and Anthropic's Claude. However, this progress is accompanied by a widening safety gap, as open-weight models lack the robust safeguards of their closed-weight counterparts.
A new report from AI safety nonprofit SaferAI highlights that GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given, a stark contrast to the consistent refusals from Anthropic's Claude Opus 4.7. This raises concerns that highly capable AI could be placed in the hands of malicious actors once the model weights are downloaded and any built-in protections are removed or modified.
While frontier developers like OpenAI and Anthropic employ measures such as classifiers and refusal training, these are ineffective against open-weight models. Henry Papadatos, executive director of SaferAI, emphasized that the objective should be to make safe capabilities accessible while removing dangerous ones, even in an open-source fashion.
Techniques like pre-training data filtering are being explored to reduce hazardous knowledge, though their practicality varies between cybersecurity and biological applications. The pressure to improve coding capabilities, a major moneymaker for AI developers, complicates efforts to limit misuse. Some developers are selectively restricting model assistance, such as Anthropic's Opus 5, which can search vulnerabilities in uncompiled but not compiled code.
Chinese leaders, including President Xi Jinping, have acknowledged the risks of advanced AI and stressed the need for human control. However, Chinese AI regulations have historically focused more on politically sensitive content than on catastrophic risks like offensive cyber capabilities. Graham Webster, a researcher at the Stanford Cyber Policy Center, noted that Chinese companies operate within a system where users are accountable, potentially offering a different approach to AI governance.
Advocates for open-weight models argue that releasing weights is crucial for cybersecurity, enabling companies to defend against attacks and prepare for future threats. Clem Delangue, CEO of Hugging Face, stated that systems used to stop AI-powered cyberattacks can also help defend against millions of daily attacks. However, Papadatos counters that this benefit is often overstated and does not justify the open-sourcing of dangerous capabilities, as attackers can adopt new tools far faster than defenders.
