Z.ai launches GLM-5.3 with claimed 50% gain on coding benchmark
Z.ai, the international brand of Chinese AI firm Zhipu, has released GLM-5.3, a coding-focused update that the company says scores 50% higher than its predecessor on an internal coding benchmark. Model weights are slated for release two weeks after launch following additi…
Intelligence analysis by Llama

Z.ai has rolled out GLM-5.3, a post-trained update to its GLM-5 series targeting coding, long-horizon agent tasks, and cybersecurity work. The company claims a 50% lift on its own Z.ai Code Bench and leading open-source scores on Terminal-Bench 3.0 and Agents' Last Exam, with open weights due in two weeks.
Z.ai is like a kid who got really good at homework by practicing a lot, not by getting a new brain. Their new helper, GLM-5.3, is better at writing computer code and finding security bugs, and they say it's 50% better on a coding test. Soon anyone will be able to download it for free, but Z.ai is waiting two weeks to double-check it's safe first.
Analysis
GLM-5.3 and the post-training story
Z.ai is positioning GLM-5.3 less as a wholesale architecture change and more as a sharpening of its existing base. The company explicitly states the new release uses the same base model as GLM-5.2, with the gains attributed to post-training. That framing matters because it signals where the lab is currently extracting value: reinforcement-style fine-tuning, curated code trajectories, and tool-use data, rather than a fresh pretraining run. For an audience tracking the cost and compute economics of frontier model work, the post-training story is also a story about marginal capability per dollar — a useful datapoint even if the headline number is vendor-reported.
50% on Z.ai Code Bench
The headline figure is a claimed 50% gain on Z.ai's own Z.ai Code Bench, paired with leading open-source results on Terminal-Bench 3.0 and Agents' Last Exam. Two of those three are external benchmarks, which raises the bar for independent verification once the weights drop. Internal benchmarks, by contrast, can be tuned to favor a release's specific training mix, so the 50% figure should be read as a company-stated delta rather than an independently audited one. The real test will be how the open weights behave on community harnesses like SWE-bench and LiveCodeBench once researchers get their hands on them in roughly two weeks.
Two-week weights and the cybersecurity angle
Alongside coding, Z.ai is leaning into cybersecurity, claiming stronger vulnerability-discovery and exploitation performance. That focus is commercially interesting for red-teaming and defensive tooling, but it also explains why the weights are not shipping on day one. The company says it will hold the open release for two weeks to complete additional safety evaluation — a step that resembles the staggered rollouts US labs have used for dual-use capabilities. The episode crystallizes a recurring tension in 2026: the more capable open-weight coding and security models get, the more carefully their publishers, regulators, and downstream developers have to choreograph release timing.
Key points
- Z.ai released GLM-5.3, an update built on the same base as GLM-5.2 with gains attributed to post-training.
- The company claims a 50% improvement on its internal Z.ai Code Bench versus GLM-5.2.
- Z.ai reports leading open-source results on Terminal-Bench 3.0 and Agents' Last Exam.
- The release emphasizes coding, long-horizon tasks, and cybersecurity, including vulnerability discovery and exploitation.
- Model weights are scheduled for open release two weeks after launch, pending additional safety evaluation.
If the claimed coding and long-horizon agent gains hold up under community testing, GLM-5.3 would give open-source developers a stronger free alternative for software engineering and security workflows. The staggered two-week safety window could also become a template for releasing dual-use capabilities in a way that satisfies both openness advocates and risk-conscious regulators.
The 50% figure comes from Z.ai's own benchmark, so independent reproduction on community suites is essential before treating the jump as real. The stronger vulnerability-discovery and exploitation performance, in particular, raises familiar dual-use concerns: even with a two-week safety review, open-weight release of an improved offensive-security model could expand access for attackers.


