Chinese AI startup Z.ai has released GLM-5.3, a model that delivers substantial improvements in long-horizon coding and agentic engineering without retraining the base model — achieving gains purely through scaled post-training. The release also surfaces an uncomfortable truth for frontier AI labs: the same capabilities that make models better at software engineering also make them more capable at security research.
Post-training at scale
GLM-5.3 builds on the 743-billion-parameter base model behind GLM-5.2, rather than introducing a new foundation. Z.ai expanded its post-training system around long-horizon reinforcement learning, using environments that resemble complete engineering jobs rather than isolated programming exercises.
The approach produced notable benchmark improvements. GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 26.2 to 48.2 on AutomationBench.
But Z.ai is emphasizing efficiency alongside performance. On its private Z.ai Code Bench, GLM-5.3 reaches 34.5% at its Max reasoning setting while consuming roughly 75,000 output tokens per task — compared with 23.4% at 96,000 tokens for GLM-5.2.
The cybersecurity wrinkle
The more consequential development is cybersecurity. Z.ai introduced vulnerability-discovery environments into GLM-5.3’s post-training expecting incremental improvement. Instead, capability progressed further along the exploitation chain than anticipated.
On CyberGym, which tests vulnerability discovery against source code, GLM-5.3 scores 84.5%, edging past GPT-5.6 Sol at 83.6% and Mythos 5 at 83.8%. On ExploitBench, GLM-5.3 scores 54.4%, more than twice GLM-5.2’s 24.4% — though well behind GPT-5.6 Sol’s 76.5%.
Already, GLM-5.3’s cyber capabilities have found what Z.ai calls a “potentially serious vulnerability” in Cursor, the AI coding assistant recently acquired by SpaceX.
The access dilemma
Z.ai says work with security teams in China has resulted in 2,436 vulnerability findings across 269 projects, with 1,097 classified as critical or high severity. The company plans to release open weights approximately two weeks after launch, “once safety evaluation and hardening are complete.”
The release is initially available only through Z.ai’s GLM Coding Plan and ZCode environment, with API access and open weights following later.
The episode illustrates a growing tension across the frontier model space: the same long-horizon agent capabilities that make models useful for engineering also make them capable security researchers — and potentially capable offensive operators. How to distribute those capabilities responsibly may become as important as how to train them.