Zhipu launches flagship model GLM-5.3 as China seeks Mythos-level edge in cyber defence
Chinese AI firm Zhipu has unveiled its flagship GLM-5.3 model, claiming it outperformed leading US systems like Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol in a key cybersecurity test, CyberGym 2.
Intelligence analysis by Gemini 2.5 Flash

Zhipu's new GLM-5.3 model marks a significant step in China's pursuit of advanced AI capabilities for cyber defense, demonstrating superior performance in identifying security flaws. However, the model lagged behind its Western counterparts in its ability to exploit those vulnerabilities, highlighting both progress and areas for further development in the competitive global AI landscape.
Imagine a super-smart computer brain that's like a detective for computer programs. China just made a new one called GLM-5.3 that's really good at finding hidden traps and weak spots in computer code, even better than some famous ones from other countries at this specific job. But, it's not as good at actually using those traps to get into a system, like a detective who finds a secret door but can't quite figure out how to open it yet.
Analysis
Zhipu, a prominent Chinese artificial intelligence firm also known as Z.ai, has introduced its latest flagship model, GLM-5.3, positioning it as a key player in China's strategic efforts to bolster its cyber defense capabilities. The launch underscores a broader national ambition to rival and potentially surpass Western advancements in AI, particularly in areas critical for national security and technological sovereignty. The company's claims of GLM-5.3 outperforming established US models in specific cybersecurity benchmarks highlight the rapid pace of AI development within China and its focused investment in this sector.
GLM-5.3
Zhipu's GLM-5.3 is presented as a frontier model designed to enhance cybersecurity measures. The company asserts that this model has undergone rigorous testing, including evaluations with Chinese security teams against real-world codebases. These tests reportedly led to the identification of 2,436 vulnerabilities across 269 projects, with 1,097 of these flaws being rated as medium to high severity after expert review. This practical application suggests a tangible impact on improving software security by proactively detecting weaknesses before they can be exploited. The model's development is part of a larger narrative of China seeking to establish a 'Mythos-level' edge in cyber defense, indicating a desire to achieve a benchmark of excellence comparable to leading international systems.
CyberGym 2
One of the key benchmarks where GLM-5.3 reportedly excelled is CyberGym 2, a test designed to measure an AI model's proficiency in identifying and validating security flaws within source code. According to Zhipu, GLM-5.3 achieved an impressive success rate of 84.5 per cent on CyberGym 2. This score notably surpassed the performance of Anthropic’s Mythos 5, which scored 83.8 per cent, and OpenAI’s GPT-5.6 Sol, which achieved 83.6 per cent. This specific triumph on CyberGym 2 positions GLM-5.3 as a leading model in the crucial task of vulnerability detection, suggesting that Chinese AI research is making significant strides in foundational cybersecurity analysis.
ExploitBench
Despite its strong performance on CyberGym 2, GLM-5.3 demonstrated a comparative weakness on ExploitBench, another critical cybersecurity benchmark. ExploitBench evaluates an AI model's ability to 'climb the exploitation ladder,' essentially measuring its capacity to not just find vulnerabilities but also to understand and leverage them for exploitation. On this benchmark, GLM-5.3 scored 54.4 per cent, which significantly trailed Anthropic’s Mythos 5 at 78 per cent and OpenAI’s GPT-5.6 Sol at 76.5 per cent. This disparity indicates that while GLM-5.3 is adept at identifying flaws, its capabilities in understanding and executing the exploitation process are not yet on par with its Western counterparts. This gap highlights a crucial area for further development if China aims for a comprehensive 'Mythos-level' edge across all facets of cyber defense and offense.
Key points
- Chinese AI firm Zhipu launched its flagship GLM-5.3 model.
- GLM-5.3 achieved an 84.5% success rate on CyberGym 2, outperforming Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol in identifying security flaws.
- The model lagged on ExploitBench, scoring 54.4% compared to Mythos' 78% and GPT-5.6 Sol's 76.5% in exploitation capabilities.
- Zhipu tested GLM-5.3 with security teams, identifying 2,436 vulnerabilities across 269 projects, with 1,097 rated medium to high severity.
- The development is part of China's broader effort to gain an edge in AI-driven cyber defense.
The advancements demonstrated by GLM-5.3 could significantly enhance China's cyber defense capabilities, leading to more secure digital infrastructure and a reduction in exploitable vulnerabilities. This progress could also foster greater global collaboration in AI development, as Zhipu itself stated that "AI development should not be a solo performance by one nation, but a symphony of global collaboration."
While strong in detection, GLM-5.3's lagging performance in exploitation could mean that identified vulnerabilities might still be exploited by more sophisticated AI systems. This disparity could also fuel an escalating AI arms race in cybersecurity, where nations focus on both defensive and offensive AI capabilities, potentially increasing global cyber risks.



