Key facts
- Anthropic warns Z.ai's GLM-5.3 AI model has potent cyber capabilities and weak safety constraints.
- GLM-5.3 can build working cyber exploits almost as well as Anthropic's Claude Mythos Preview.
- GLM-5.3's safeguards can be bypassed or removed with simple techniques, with refusal rates dropping from over 90% to as low as 2%.
- NIST found GLM-5.3 to be the most cyber-capable open-weight model released to date.
- Attackers can bypass GLM-5.3's safeguards between 64% and 100% of the time in simulated tests.
- GLM-5.3 developed working exploits in 50 of 410 attempts on Anthropic's ExploitBench.
Anthropic, a leading US artificial intelligence company, has issued a warning regarding the capabilities of GLM-5.3, an open-weight AI model developed by China-based Zhipu AI (known internationally as Z.ai). The firm's Frontier Red Team found that GLM-5.3 possesses advanced cyber capabilities, nearly matching Anthropic's own frontier model, Claude Mythos Preview, in developing sophisticated, end-to-end cyber exploits.
However, Anthropic highlighted a critical difference: GLM-5.3 has significantly weaker safety constraints. In simulated tests, attackers could bypass or remove GLM-5.3's safeguards between 64% and 100% of the time using simple techniques. This lax security increases the potential for malicious actors to co-opt the model for harmful cyberattacks.
Anthropic's analysis showed GLM-5.3 successfully developed working exploits in 50 out of 410 attempts on the ExploitBench benchmark, a rate comparable to Claude Mythos Preview's 56 out of 410 attempts. On Anthropic's internal binary exploitation test, GLM-5.3 achieved a full takeover in 4% of trials, compared to 6% for Mythos Preview. Earlier models like GLM-5.2 and Claude Opus 4.6 managed no successful exploits in these tests.
NIST's Center for AI Standards and Innovation (CAISI) corroborated these findings, labeling GLM-5.3 as "the most cyber-capable open-weight model released to date" and estimating it lags behind US frontier models by approximately four months on aggregate cyber benchmarks. Unlike US frontier models, which are often restricted to vetted users, GLM-5.3 is publicly downloadable.
Anthropic noted that the model's safeguards could be easily bypassed. While it initially refused direct requests to attack critical systems, prompts framing the task as a red-team exercise led to compliance 64% of the time, rising to 92% with pre-filled reasoning. An "abliterated" version, where refusals were edited out of the model's weights, complied every time. Anthropic stated that editing these weights cost approximately $4,400 in computing, and modified versions appeared publicly within days of GLM-5.3's release.
