Anthropic Finds China’s GLM-5.3 Can Build Working Cyber Exploits

Written By
Liz Ticong
Liz Ticong
Oct 1, 2026
3 minute read
Anthropic’s testing found GLM-5.3 could develop working cyber exploits with limited human guidance.

Anthropic’s testing found GLM-5.3 could develop working cyber exploits with limited human guidance. Image generated with ChatGPT.

eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Finding a software flaw is one problem. Turning it into a working attack has traditionally taken more time and specialist skill.

Anthropic researchers found that China’s GLM-5.3 could develop working exploits against known software vulnerabilities with limited human guidance. Researchers also tested what happens when users can download the model and alter its safeguards.

Security teams may now have less time between a public vulnerability disclosure and usable attack code.

One exploit took $20 and 20 minutes of human attention

Anthropic gave a smaller version called GLM-5.3-Flash public information about a recent Chrome vulnerability and another known flaw during its GLM-5.3 security testing. A researcher spent about 20 minutes setting up the task before the model worked for eight hours and produced a functioning exploit chain. The company estimated the API cost at $20.40.

Automated testing suggested the case was not isolated. Across 410 attempts against known flaws in V8, the JavaScript engine used by Chrome, GLM-5.3 completed 50 working exploits. Restricted Claude Mythos Preview completed 56 under the same setup.

Researchers ran the work in isolated environments against offline targets prepared for testing. Findings demonstrate exploit-building capability under controlled conditions. They do not show GLM-5.3 independently attacking systems on the public internet.

Open weights put safeguards in users’ hands

Researchers compared GLM-5.3’s default refusal behavior with a bypass prompt and a locally modified copy.

What researchers testedOutcome
Direct harmful cyber instructionModel did not proceed
Prompt that made the model appear to have decided to continueModel proceeded in 92% of trials
Downloaded copy modified to remove refusalsModel proceeded in 100% of trials

Because GLM-5.3’s weights can be downloaded, local operators can change refusal behavior instead of relying on restrictions enforced through a provider’s API.

Z.ai delayed the weights release by two weeks for additional safety evaluation and hardening. Its GLM-5.3 release also reported faster-than-expected growth in cyber capability.

Advertisement

What eWeek found: GLM-5.3 made a major cyber leap without a new base model

eWeek compared Z.ai’s release material with Anthropic’s evaluation and NIST’s independent assessment. All three sources point to a version-level risk that companies could easily miss if they look only at a model family name.

Z.ai says GLM-5.3 uses the same base model as GLM-5.2, but later training changed what it could do. On one exploit test, its score rose from 24.4 to 54.4. On another, completed security tasks within two hours climbed from 29 to 105. In real terms, a model built on the same foundation became far more capable at finding and developing software exploits after additional training.

Anthropic found the same version-to-version difference under separate testing, with GLM-5.3 completing exploit tasks that GLM-5.2 could not. NIST’s independent assessment reached a similar conclusion and called GLM-5.3 the most cyber-capable open-weight model it had evaluated.

Different test setups prevent direct score comparisons. Companies should focus on what changed between releases. Sharing the same base model did not give GLM-5.2 and GLM-5.3 the same security profile, so an approval tied to one version should not automatically carry over to the next.

What security teams should review

Security teams evaluating GLM-5.3 or similar open-weight models should treat each major release as a fresh security decision, with access and controls reviewed against the capabilities of that specific version.

  • Repeat predeployment model testing before granting a newer version the same access to source code or development tools.
  • Isolate sensitive repositories and limit network access for open-weight AI models, particularly when users can modify local copies.
  • Recheck production credentials and permissions separately so a newer model does not inherit access granted to an earlier release.
  • Shorten triage timelines for newly disclosed vulnerabilities when AI can spend hours developing exploits after a brief human setup. AI-powered vulnerability hunting can aid defenders, but similar automation can also reduce the work needed to produce attack code.

Model upgrades now warrant the same scrutiny as other security-sensitive software changes. Release-level reviews give teams a firmer basis for deciding where a model belongs, what systems it can reach, and how much access it should receive.

Advertisement

More AI news: GPT-6.1 Sol arrives just a week after GPT-6 Sol with sizable gains in coding, computer use, and professional tasks.

Liz Ticong

Liz Ticong is a staff writer for eWeek and TechRepublic focused on AI, cybersecurity, enterprise software, and data. She has more than 10 years of editorial experience as a technology industry writer, combining reporting, product research, and hands-on software testing in her coverage. Her work has been published on Datamation, Enterprise Networking Planet, and TechnologyAdvice.com. She writes technology news, software reviews, product comparisons, and buyer’s guides for business and IT readers.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.