Category: AI Safety • Cybersecurity

What it is
Z.ai released GLM-5.3, built on the same 743-billion-parameter base as GLM-5.2 with gains from post-training alone, and reported it scored 84.5% on CyberGym (vulnerability detection), narrowly ahead of Mythos 5’s 83.8% and GPT-5.6 Sol’s 83.6%, per Z.ai’s own figures. The gap widens sharply on turning flaws into working attacks: GLM-5.3 scored 54.4% on ExploitBench versus Mythos 5’s 78%, and completed fewer tasks on time-bound ExploitGym tests. Z.ai is delaying public weights roughly two weeks for safety hardening, calling it the first time a Chinese lab has cited emergent offensive capability as a reason to hold back a release.
Why it Matters for Enterprises
Open-weight models are approaching frontier labs on vulnerability discovery even as they trail on exploitation, narrowing the defensive-tooling gap while raising dual-use risk. Security teams should track open-model capability alongside access controls on restricted models like Mythos.