Z.ai says its open-source GLM-5.3 model can identify software vulnerabilities at a level close to, and slightly ahead of, Anthropic’s restricted Mythos 5 model on one reported benchmark. The Chinese AI startup reported an 84.5% score on CyberGym, compared with 83.8% for Mythos 5.
Those figures have not been independently verified. Still, the announcement places GLM-5.3 among the latest Chinese AI models being presented as competitors to leading U.S. systems in cybersecurity and coding work.
GLM-5.3 Vulnerability Detection Performance
CyberGym measures vulnerability-detection performance. According to Z.ai’s reported results, GLM-5.3 achieved a higher score than Mythos 5 when identifying software flaws.
|
Benchmark
|
GLM-5.3
|
Mythos 5
|
|
CyberGym vulnerability detection
|
84.5%
|
83.8%
|
The difference is narrow, and the results remain unverified. Z.ai’s claim is specifically about finding vulnerabilities; it does not indicate that the model performed better across every cybersecurity task.
GLM-5.3 Trailed Mythos 5 on Exploit Development
Z.ai’s own results show a different picture when the task moves from detecting a flaw to turning that finding into a working exploit. On ExploitBench, GLM-5.3 scored 54.4%, while Mythos 5 scored 78.0%.
|
Benchmark
|
GLM-5.3
|
Mythos 5
|
|
ExploitBench
|
54.4%
|
78.0%
|
In a timed test, GLM-5.3 completed 105 attack-development tasks in two hours. Anthropic’s model completed 181 tasks over the same period.
Z.ai’s reported benchmark results therefore suggest that GLM-5.3 was stronger at vulnerability detection than at exploit development when measured against Mythos 5.
Coding and Agent Benchmark Scores
GLM-5.3 has 743 billion parameters. Beyond cybersecurity testing, Zhipu AI, the company behind the Z.ai brand, reported results in coding and agent benchmarks.
The model received a 28.3 score on Terminal-Bench 3.0, which Zhipu said was the highest result among open-source models. It was reported to have ranked ahead of Moonshot AI’s Kimi K3 on that benchmark.
Zhipu also reported a 66.9 score for GLM-5.3 on DeepSWE 1.1. The company said the model’s coding and agent capabilities approach those of Anthropic’s Claude Fable 5.
These results are company-reported figures, rather than independently verified comparisons.
Open-Source Access and Restricted Cybersecurity Tools
Z.ai framed GLM-5.3 as a challenge to the closed-access approach used for Mythos. According to the reported description, Mythos is a version of Claude Fable 5 with cybersecurity safeguards removed and is available only to vetted organizations.
Z.ai argued that advanced cyber-defense tools should be available to open-source developers and smaller security teams, rather than being controlled by a limited number of closed-model providers.
The company said it plans to publicly release GLM-5.3 after completing safety evaluations. However, it said the model’s most sensitive cybersecurity capabilities would be available through a “trusted access” program for verified users.
Z.ai also announced an Open Source Shield initiative. The initiative is intended to audit open-source projects and provide access to defensive models.
GLM-5.3 Safety Protections
Public release of model weights can make safeguards more difficult to enforce because users may alter models or combine them with external tools. That concern remains relevant for AI systems with cybersecurity capabilities.
Z.ai said it added several protections to GLM-5.3, including:
- Systems designed to screen risky requests
- Monitoring of model outputs
- Training intended to make the model reject malicious tasks
The company’s planned trusted-access approach for sensitive cybersecurity functions adds another layer to its stated release plan.
Other Chinese Claims of Mythos-Level Cybersecurity AI
Z.ai is not the first Chinese company to claim capabilities comparable to Mythos. Cybersecurity company 360 said its Tulongfeng system achieved equivalent performance by combining AI models, security data, and automated tools.
Those claims were also unverified.
GLM-5.3 follows Z.ai’s earlier GLM-5.2 model, which gained popularity with overseas developers. The earlier model was described as offering capabilities that approached leading U.S. models at a lower cost.

