Beijing-based AI firm Z.ai on Friday unveiled GLM-5.3, an open-weights model that claims top-tier programming benchmarks and competitive vulnerability-detection scores against leading U.S. systems from Anthropic and OpenAI.
Built on the same 700-billion-parameter base model as its June predecessor, GLM-5.2, the company said all performance gains stem from expanded reinforcement learning and synthetic post-training environments.
Public release of the model's weights will be delayed for two weeks while safety evaluations and security hardening are completed, according to the company.
Benchmark gains and cybersecurity leaps
Z.ai reported that GLM-5.3 delivers a 50% improvement over GLM-5.2 on its internal Z.ai Code Bench. On public evaluations, the system reached 88.2 on Terminal Bench 2.1, 28.3 on Terminal Bench 3.0, and 66.9 on DeepSWE v1.1.
The company highlighted emergent cybersecurity capabilities that evolved during post-training. On the CyberGym vulnerability-detection benchmark, GLM-5.3 scored 84.5%, edging out Anthropic's Mythos 5 at 83.8% and OpenAI's GPT-5.6 Sol at 83.6%.
However, the model lags behind U.S. closed-frontier systems in building active exploits. On ExploitBench, GLM-5.3 scored 54.4%, trailing Mythos 5's 78.0%. In a six-hour timed test on ExploitGym, GLM-5.3 completed 130 attack-development tasks compared to 247 tasks logged by Mythos 5.
Alongside the announcement, Z.ai launched the Z.ai Security Disclosure Ledger, documenting 2,436 vulnerabilities identified across 269 open-source software projects, including Linux kernel components and Apache projects.
Controlled access and safety measures
To manage offensive risks, Z.ai said it is implementing a tiered rollout that keeps sensitive exploit features restricted to vetted security partners under a trusted access program.
The company stated on X that it added safeguards to reject malicious requests while supporting legitimate defensive testing, writing: "An open world cannot have only open attack surfaces. It must also have an open shield."
"To the best of my knowledge, this is the first time a Chinese lab is publicly justifying a delayed open release of model weights with safety considerations," said Gabriel Wagner, an AI governance researcher at Concordia AI, per Reuters.
Strong technical results have not erased questions about Z.ai’s commercial position.
Despite the product claims, shares of Hong Kong-listed Z.ai fell nearly 4% Friday. Bloomberg Intelligence analyst Robert Lea noted that "This firm remains on a completely unsustainable commercial footing," adding that "Rising agentic AI will drive Z.ai’s inference costs and losses higher."
What the launch means
GLM-5.3’s significance is less about whether it has overtaken every leading U.S. model than about how competitive Chinese systems are becoming in specialized coding and cybersecurity tasks.
Z.ai is showing that stronger coding models can emerge through post-training without replacing the underlying model. At the same time, its cybersecurity results suggest that general-purpose coding systems can develop offensive capabilities faster than expected as training becomes more sophisticated.
The more difficult question is what happens when those improvements reach security-sensitive tasks. Z.ai is delaying the open release of GLM-5.3’s weights while it completes safety work, underscoring the tension between open access and the risk that stronger defensive capabilities can also be repurposed for attack.
Other News: DeepSeek’s V4 Pro is challenging Anthropic’s Claude Opus 5 for developers, intensifying competition over which AI models perform best on coding tasks.


