What Zhipu’s personal GLM-5.3 information says about the benchmark hole


Zhipu’s release note for GLM-5.3 incorporates a sentence that did not make it into most of the protection. Describing its personal cybersecurity outcomes, the Beijing firm writes that functionality “is rising quickest precisely the place we are furthest behind.”

Zhipu, which additionally trades as Z.ai, is considered one of a handful of Chinese language labs releasing fashions that compete with the American frontier. On August 14, it launched GLM-5.3, a coding-focused mannequin, and printed a technical launch observe setting out how the mannequin performs towards its rivals. That observe is the supply for the whole lot reported right here.

The declare that travelled was about safety. Alongside the coding outcomes, Zhipu mentioned GLM-5.3 had grow to be unexpectedly good at discovering software program vulnerabilities, scoring 84.5% on a benchmark referred to as CyberGym towards 83.8% for Anthropic’s Mythos 5 and 83.6% for OpenAI’s GPT-5.6 Sol. Headlines adopted reporting {that a} Chinese language mannequin now out-finds the American ones at bug searching.

The rationale that lands tougher than a common benchmark outcome is what vulnerability discovery has grow to be. A mannequin that may learn a codebase and find exploitable flaws is helpful to a defender auditing their very own software program and helpful to anybody doing the similar to anyone else’s. Anthropic’s equal work sits behind restricted entry for that purpose, whereas Zhipu intends to publish GLM-5.3’s weights for anybody to obtain.

Zhipu’s personal launch is extra measured than the protection it produced. The CyberGym quantity is actual, and it is in the paper. It is additionally the narrowest of the three cybersecurity outcomes the firm printed, and Zhipu is upfront that the different two go the different approach.

Three benchmarks, three totally different photos

CyberGym begins from supply code the mannequin can learn and assessments whether or not it could discover a vulnerability and make sure the flaw is real. That is the outcome that travelled, and the margin is seven tenths of a share level.

ExploitBench asks one thing tougher, requiring the mannequin to purpose about an actual vulnerability and the way it might be exploited. GLM-5.3 scores 54.4%, greater than double its predecessor’s 24.4%. Mythos 5 scores 78.0% and GPT-5.6 Sol 76.5%.

ExploitGym counts what number of exploitation duties a mannequin finishes inside a set time finances. GLM-5.3 completes 105 duties in two hours and 130 in six. Mythos 5 completes 181 and 247.

These two outcomes have been reported thinly, and so they are the ones that describe the hole. Discovering a flaw and constructing a working exploit from it are totally different jobs. Zhipu’s studying is that the additional alongside that chain a check sits, the additional behind its mannequin is, and the firm says so in the launch slightly than leaving it to be found.

Bar chart comparing five AI models across three cybersecurity benchmarks. GLM-5.3 leads on CyberGym at 84.5 but trails Mythos 5 on ExploitBench and ExploitGym.
Zhipu’s personal comparability throughout the three cybersecurity benchmarks. GLM-5.3 leads on CyberGym, the vulnerability discovery check, and falls behind Anthropic’s Mythos 5 on each exploitation measures. Supply: Z.ai.

Which Anthropic mannequin, and why it retains altering

A part of the confusion in the protection comes from Zhipu evaluating three totally different Anthropic fashions in three totally different locations. The principle benchmark desk units GLM-5.3 towards Opus 4.8. The efficiency charts use Fable 5. The cybersecurity part makes use of Mythos 5. Anybody studying rapidly comes away with a single comparability that does not exist.

On coding, the image is blended slightly than dominant. GLM-5.3 leads Opus 4.8 on some assessments and trails it on others, and Zhipu states plainly that its mannequin stays behind Claude Fable 5 on the firm’s personal inside coding benchmark.

How the assessments have been run

The methodology footnotes comprise one thing the summaries skipped. Zhipu evaluated GLM-5.3 on CyberGym, ExploitGym, ExploitBench, Terminal Bench and several other different duties inside Claude Code 2.1.207, Anthropic’s coding agent.

That is not improper. Utilizing a standard harness throughout fashions is how a comparability stays truthful, and Zhipu paperwork the settings it used. It is value noticing anyway. A Chinese language open-weights mannequin’s frontier claims are being measured via American agent software program, which says one thing about the place the tooling layer sits on this competitors that the mannequin scores do not.

Two additional details deserve consideration before the CyberGym outcome is handled as settled. The rating is a single run, reported as move@1 throughout 1,507 duties, with no variance figures given. A niche of seven tenths of a degree between two single runs is not a niche anybody ought to lean on. And the ExploitGym time budgets have been normalised utilizing throughput charges from Synthetic Evaluation, with rescaling elements listed for GLM-5.3, Kimi K3 and Qwen3.8 Max, however not for Mythos 5.

The vulnerability rely and the quantity that is lacking

Past the benchmarks, Zhipu says it labored with safety groups in China to run its fashions towards actual codebases, figuring out 2,436 vulnerabilities throughout 269 open-source tasks. The severity break up is 107 essential, 990 excessive, 1,286 medium and 53 low. The oldest flaw dates to 1981, and the common vulnerability had been sitting in code for 26.6 years before it was discovered.

Summary panel from Zhipu's release showing 2,436 findings tracked, 53 publicly disclosed, 2,383 under embargo and 1,097 labelled critical and high.
Zhipu’s disclosure abstract. The panel labels 1,097 findings as essential and excessive, matching the severity breakdown of 107 essential and 990 excessive. The physique textual content of the similar launch describes the determine as medium-to-high. Supply: Z.ai.

One discrepancy is value carrying rigorously. Zhipu’s abstract panel labels 1,097 findings as essential and excessive, which matches the severity desk. The physique textual content of the similar launch describes these 1,097 as medium-to-high. A number of retailers have reproduced the second model.

The rely additionally arrives after what Zhipu describes as skilled evaluate, screening and deduplication, so the uncooked mannequin output is not what is being reported. Of the 2,436 findings, 53 have been publicly disclosed, and a pair of,383 stay beneath embargo. The discharge does not say what number of have been beforehand unknown, and it does not say what number of have been independently reproduced. These are the two figures that may flip a quantity declare right into a functionality declare.

What issues greater than the benchmark desk

Two issues in the launch have longer penalties than the CyberGym margin.

The primary is effectivity. Zhipu stories GLM-5.3 reaching 31.4% on its inside coding benchmark at round 50,000 output tokens per process, towards Opus 4.8 at 29.5% utilizing 120,000. Barely higher work for lower than half the tokens is a price argument, and value determines whether or not safety groups outdoors the largest budgets can run these instruments in any respect.

The second is distribution. Zhipu says the weights can be printed as soon as security analysis and hardening are completed. That has not occurred but, and till it does, the open-weights declare is a dedication slightly than a reality. If it holds, a mannequin with documented vulnerability-discovery functionality turns into one thing any group can obtain and run regionally, together with in markets that can by no means have entry to an export-controlled American mannequin.

The weights are due at the finish of August.

See additionally: Anthropic walks into the White House and Mythos is the reason Washington let it in

Banner for the AI & Big Data Expo event series.

Need to study extra about AI and massive information from trade leaders? Take a look at AI & Big Data Expo going down in Amsterdam, California, and London. The great occasion is a part of TechEx and is co-located with different main expertise occasions together with the Cyber Security & Cloud Expo. Click on here for extra information.

AI Information is powered by TechForge Media. Discover different upcoming enterprise expertise occasions and webinars here.




Disclaimer: This article is sourced from external platforms. OverBeta has not independently verified the information. Readers are advised to verify details before relying on them.

0
Show Comments (0) Hide Comments (0)
0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments

Stay Updated!

Subscribe to get the latest blog posts, news, and updates delivered straight to your inbox.