All NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
All NewsHome
← Back to AI & Technology

Open-weight AI models approach frontier capabilities, but safety gap widens

Created at 4 Aug · 8:11 PM1 source↑ Market-relevant
IN SHORT

A Chinese open-weight AI model, GLM-5.2, is closing the gap with leading systems like GPT-5.5 and Claude Opus 4.7 in capabilities, according to SaferAI. However, the report highlights a growing disparity between advanced AI functionalities and robust safety practices, particularly concerning the potential for misuse of open-weight models.

Key Numbers

5.6version of OpenAI's GPT model
5.2version of Z.ai's GLM model
4.7version of Anthropic's Claude Opus model

Who's Involved

SaferAI
AI safety nonprofit that evaluated GLM-5.2
Z.ai
Chinese company that developed GLM-5.2
Henry Papadatos
Executive Director of SaferAI
OpenAI
Developer of GPT models
Anthropic
Developer of Claude models
Xi Jinping
Chinese President
Graham Webster
Chinese AI policy researcher at Stanford Cyber Policy Center
Clem Delangue
CEO of Hugging Face
Open-weight AI models approach frontier capabilities, but safety gap widens

↳ Why This Matters

The rapid advancement of open-weight AI models, while offering potential benefits for cybersecurity and accessibility, poses significant risks due to their lack of robust safety measures. This development challenges policymakers and developers to find effective ways to manage the potential misuse of powerful AI technologies, balancing innovation with global security.

Key facts

  • GLM-5.2, an open-weight AI model from China's Z.ai, is reportedly only a few months behind industry leaders in cyber and bio capabilities.
  • SaferAI's evaluation found GLM-5.2 refused no offensive cyber or dual-use biology tasks, while Claude Opus 4.7 consistently refused.
  • Open-weight models, once downloaded, can have their safeguards removed or modified, making them difficult to police.
  • Frontier AI developers like OpenAI and Anthropic use safeguards that are ineffective on open-weight models.
  • Chinese President Xi Jinping has called for strict human control over AI, acknowledging risks associated with advanced AI.
  • Graham Webster notes that Chinese AI regulations have historically prioritized politically sensitive content over catastrophic AI risks.

Open-weight AI models are rapidly advancing in capability, with China's Z.ai GLM-5.2 reportedly narrowing the gap with industry leaders like OpenAI's GPT and Anthropic's Claude. However, this progress is accompanied by a widening safety gap, as open-weight models lack the robust safeguards of their closed-weight counterparts.

A new report from AI safety nonprofit SaferAI highlights that GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given, a stark contrast to the consistent refusals from Anthropic's Claude Opus 4.7. This raises concerns that highly capable AI could be placed in the hands of malicious actors once the model weights are downloaded and any built-in protections are removed or modified.

While frontier developers like OpenAI and Anthropic employ measures such as classifiers and refusal training, these are ineffective against open-weight models. Henry Papadatos, executive director of SaferAI, emphasized that the objective should be to make safe capabilities accessible while removing dangerous ones, even in an open-source fashion.

Techniques like pre-training data filtering are being explored to reduce hazardous knowledge, though their practicality varies between cybersecurity and biological applications. The pressure to improve coding capabilities, a major moneymaker for AI developers, complicates efforts to limit misuse. Some developers are selectively restricting model assistance, such as Anthropic's Opus 5, which can search vulnerabilities in uncompiled but not compiled code.

Chinese leaders, including President Xi Jinping, have acknowledged the risks of advanced AI and stressed the need for human control. However, Chinese AI regulations have historically focused more on politically sensitive content than on catastrophic risks like offensive cyber capabilities. Graham Webster, a researcher at the Stanford Cyber Policy Center, noted that Chinese companies operate within a system where users are accountable, potentially offering a different approach to AI governance.

Advocates for open-weight models argue that releasing weights is crucial for cybersecurity, enabling companies to defend against attacks and prepare for future threats. Clem Delangue, CEO of Hugging Face, stated that systems used to stop AI-powered cyberattacks can also help defend against millions of daily attacks. However, Papadatos counters that this benefit is often overstated and does not justify the open-sourcing of dangerous capabilities, as attackers can adopt new tools far faster than defenders.

Frequently asked questions

An open-weight AI model is one where the model's weights, which are the parameters learned during training, are publicly released. This allows anyone to download, run, and modify the model on their own hardware.

The primary concern is that open-weight models can be used by malicious actors for harmful purposes, such as cyberattacks or biological engineering, because their built-in safety safeguards can be easily removed or bypassed once the weights are downloaded.

Frontier AI developers like OpenAI and Anthropic typically rely on safeguards such as classifiers, refusal training, and API-level controls to limit dangerous outputs. However, these measures are often bypassed through 'jailbreaks'.

CyberGym is a benchmark used to evaluate the cybersecurity capabilities of AI models.

What Happens Next

01Z.ai has been asked whether it conducted internal or third-party frontier safety evaluations before releasing GLM-5.2.
02Policymakers continue to debate how to govern increasingly powerful AI systems.

How It Developed

A Chinese open-weight AI model, GLM-5.2, is nearing the capabilities of leading AI systems like OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7.
A report by SaferAI indicates that GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given, contrasting with Claude Opus 4.7's consistent refusals.
The increasing capability of open-weight models raises concerns about potential attackers gaining access to powerful AI with no oversight.
Frontier AI developers typically use safeguards like classifiers and refusal training, which are ineffective on open-weight models.
Pre-training data filtering is suggested as a method to reduce hazardous knowledge, though it is more practical for biological risks than cybersecurity.
Chinese President Xi Jinping has stressed the importance of open-weight models while emphasizing strict human control over AI.
Chinese AI regulations have historically focused on politically sensitive content rather than catastrophic risks like offensive cyber capabilities.
Advocates argue open-weight models aid cybersecurity by allowing for defense against attacks and preparation for future threats.

Sources

T1
Open-weight AI models are catching up to the frontier. The safety gap remains.TechCrunch

Related Stories

Frontier AI labs lack public containment plans for rogue models, study finds
22 Aug · 4:11 PM
Mysterious Free AI Model 'Ox Alpha' Impresses Developers Amidst Origin Speculation
22 Aug · 8:56 PM
China's AI advancements offer new avenues for Asian markets
23 Aug · 8:36 AM
AI Watermark Remover Creator Overwhelmed by Viral Attention
23 Aug · 11:35 AM
Bitcoin Red Team Fights AI-Powered Exploits Amidst Model Restrictions
22 Aug · 3:35 PM