
Chinese artificial intelligence startup Z.ai says its latest GLM-5.3 model has reached a level of cybersecurity performance close to Anthropic’s restricted Mythos 5, although the two models show a substantial difference when vulnerabilities have to be turned into working attacks.
Z.ai said on August 14, 2026, that GLM-5.3 scored 84.5% on CyberGym, a benchmark that tests whether an AI system can examine software, identify vulnerabilities and confirm that the flaws are real. The company said the result was slightly higher than the 83.8% reported for Anthropic’s Mythos 5. Reuters reported that the results have not been independently verified.
The comparison changes considerably on ExploitBench, which evaluates the ability to turn discovered vulnerabilities into working exploits. Z.ai said GLM-5.3 scored 54.4%, well below Mythos 5’s 78.0% result. In a separate timed evaluation, Z.ai said GLM-5.3 completed 105 attack-development tasks in two hours and 130 in six hours, while Mythos 5 completed 181 and 247 respectively.
The figures mean GLM-5.3 has not matched Mythos 5 across the full range of cybersecurity testing. Its reported advantage is concentrated in vulnerability identification and verification, while Anthropic’s model retains a significant lead in developing working attacks from those vulnerabilities.
Z.ai is presenting GLM-5.3 as a general-purpose coding model rather than a system designed specifically for cybersecurity. Reuters reported that the company used the same base model as GLM-5.2 and developed the newer system through additional post-training and reinforcement learning in longer and more varied task environments. Z.ai says the approach produced stronger performance on complex coding and long-horizon tasks without retraining the underlying base model.
The cybersecurity results also represent a sizeable improvement over GLM-5.2 in Z.ai’s reported testing. GLM-5.2 scored 77.2% on CyberGym and 24.4% on ExploitBench, compared with 84.5% and 54.4% respectively for GLM-5.3. That represents an increase of 7.3 percentage points on CyberGym and 30 percentage points on ExploitBench.
Anthropic has treated comparable cybersecurity capabilities as a controlled-access capability. The company describes Mythos 5 as its most capable model for cybersecurity and biology research and says access has been limited to a small group of vetted partners. Anthropic has also published research showing why exploit-development capabilities have prompted tighter controls around its cybersecurity models.
Z.ai said it will also restrict some of GLM-5.3’s most sensitive cybersecurity functions. The company plans to complete additional security assessments and strengthen its safeguards before releasing the model weights publicly, with the release expected in about two weeks. Reuters reported that access to the most sensitive functions will be provided through a “trusted access” programme for verified users.
According to Z.ai, the safeguards include systems intended to screen risky requests, monitor the model’s work and train it to reject malicious tasks. The company said the controls are designed to distinguish harmful activity from legitimate uses, including fixing software bugs, cybersecurity education and authorised security testing.
The company acknowledged, however, that these protections become more difficult to enforce once model weights are available for others to download, modify or combine with external tools. Reuters reported that critics have raised concerns about maintaining those safeguards after public release.
Z.ai is also using the release to promote an initiative called Open Source Shield. The company said the programme will involve auditing selected open-source projects, providing model access for defensive security work and adding code-auditing capabilities to its ZCode programming product.
The company argues that advanced AI-assisted cybersecurity tools should be available to open-source developers and smaller security teams rather than being limited to a small number of providers of closed AI systems. The approach puts Z.ai’s planned release in direct contrast with the restricted-access model used by Anthropic for Mythos 5.
GLM-5.3’s cybersecurity performance is part of a broader rise in the capabilities of Z.ai’s GLM family. Reuters noted that the company’s previous GLM-5.2 model attracted attention among developers outside China for its coding and agent capabilities. Hugging Face also said in July that it had used GLM-5.2 while responding to a cyberattack involving a rogue OpenAI agent.
For now, the benchmark results show a more complicated picture than a simple contest between the two models. GLM-5.3 has a slightly higher reported CyberGym score than Mythos 5, but Mythos 5 remains substantially ahead on ExploitBench and on the timed attack-development tests reported by Z.ai. More broadly, the CyberGym figures are company-reported comparisons and have not been independently verified.
Z.ai’s decision to delay the open release of GLM-5.3’s weights also reflects the tension created by increasingly capable AI systems that can be used for both defensive security work and offensive activity. The company is seeking to make the model more broadly available while keeping its most sensitive cybersecurity capabilities behind additional controls.
Sources: Reuters; Anthropic Mythos 5; Anthropic exploit evaluation research; Z.ai research on GLM
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.


