GLM-5.3 is right here with superior cyber capabilities — and reportedly already discovered a ‘severe vulnerability’ in Cursor


Chinese language AI startup Z.ai, recognized internationally for its rising lineup of highly effective, largely open supply GLM sequence of language fashions, today released GLM-5.3 with substantial positive aspects in long-horizon coding and a extra consequential — and doubtlessly delicate — bounce in cybersecurity capabilities.

Already, GLM-5.3’s cyber capabilities have discovered a “doubtlessly severe vulnerability in Cursor,” the AI coding startup just lately acquired by SpaceX, in accordance to z.ai developer advocate Lou, posting on X. VentureBeat additionally tagged Cursor for affirmation on X and is awaiting response.

GLM-5.3 is obtainable initially solely by the firm’s GLM Coding Plan and ZCode coding setting, whereas API entry and open weights are coming later, “as soon as security analysis and hardening are full,” in accordance to the firm.

Z.ai says it plans to launch weights roughly two weeks after launch.

For enterprise builders, the notable a part of the launch is not merely one other spherical of benchmark enhancements. Z.ai says GLM-5.3 makes use of the similar base mannequin as GLM-5.2, with the enhancements coming fully from scaling post-training throughout extra environments, extra various duties and extra reinforcement-learning compute.

That makes GLM-5.3 one thing of a take a look at of how far a frontier-scale base mannequin will be pushed with out one other costly pretraining cycle.

“Scaling post-training is all we did for GLM-5.3,” Z.ai wrote in its technical announcement.

The outcomes counsel appreciable headroom. However they’ve additionally produced an uncommon drawback for an open-model developer: in accordance to Z.ai, cybersecurity capabilities improved sooner than anticipated as coaching scaled, notably as duties progressed from vulnerability identification towards establishing full exploitation chains.

Reuters reported Friday that Z.ai is additionally introducing controls round a few of the mannequin’s extra superior capabilities, together with a “trusted entry” method for delicate performance.

A big bounce in coding with out one other base mannequin

GLM-5.3 builds on the 743-billion-parameter-scale base mannequin behind GLM-5.2 relatively than changing it. Z.ai as an alternative expanded the post-training system it had already assembled round long-horizon reinforcement studying.

These environments more and more resemble full engineering jobs relatively than remoted programming workouts.

Z.ai describes eventualities by which an agent receives entry to codebases, documentation, compute clusters, storage methods and experimental outcomes, then has to diagnose issues, modify methods, run experiments and exhibit a measurable enchancment whereas preserving correctness. Some duties are designed to approximate a number of days of labor for an skilled engineer.

The method produced sizable generation-over-generation enhancements on Z.ai’s reported evaluations.

GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 26.2 to 48.2 on AutomationBench. On Brokers’ Final Examination CLI, it improves from 23.8 to 28.5.

The mannequin does not dominate each frontier competitor. Z.ai’s personal benchmark desk reveals GPT-5.6 Sol at 34.6 and Claude Fable 5 at 33.7 on Terminal-Bench 3.0, in contrast with GLM-5.3’s 28.3. On DeepSWE v1.1, GLM-5.3 scores 66.9, in contrast with 72.7 for GPT-5.6 Sol and 69.7 for Fable 5.

However Z.ai is additionally emphasizing effectivity relatively than benchmark place alone.

On its non-public Z.ai Code Bench, GLM-5.3 reaches a 34.5% consequence at its Max reasoning setting whereas consuming roughly 75,000 output tokens per job. GLM-5.2 reaches 23.4% whereas consuming roughly 96,000. At Excessive effort, GLM-5.3 reaches 31.4% at roughly 50,000 output tokens, in contrast with Z.ai’s reported 29.5% for Claude Opus 4.8 utilizing 120,000.

As a result of Code Bench is Z.ai’s personal non-public analysis, these comparisons must be handled as company-reported outcomes relatively than impartial measurements. Nonetheless, decreasing token consumption whereas bettering job completion is operationally necessary for enterprises deploying coding brokers, the place long-running loops could make inference value and latency compound rapidly.

Cyber capabilities developed sooner than Z.ai anticipated

The extra uncommon improvement is cybersecurity.

Z.ai launched vulnerability-discovery environments into GLM-5.3’s post-training combine anticipating the mannequin to enhance at discovering software program flaws. As an alternative, the firm says functionality started progressing additional alongside the exploitation chain.

“As we scaled post-training, cyber functionality developed sooner than we anticipated,” Z.ai wrote.

z.ai benchmarks for GLM-5.3

z.ai benchmarks for GLM-5.3. Credit score: z.ai

On CyberGym, which checks vulnerability discovery and validation towards supply code, GLM-5.3 scores 84.5%, in contrast with 77.2% for GLM-5.2. That additionally edges Z.ai’s reported scores for GPT-5.6 Sol at 83.6% and Mythos 5 at 83.8%.

The benefit does not lengthen throughout the total exploitation stack. GLM-5.3 scores 54.4% on ExploitBench, greater than twice GLM-5.2’s 24.4%, however stays nicely behind the 76.5% Z.ai reviews for GPT-5.6 Sol and 78% for Mythos 5.

Equally, on ExploitGym, GLM-5.3 completes 105 duties beneath a normalized two-hour price range and 130 beneath six hours, up from 29 and 39 for GLM-5.2. Fable 5 reaches 181 and 247, whereas GPT-5.6 Sol reaches 216 and 293.

The path of journey might matter greater than the leaderboard place.

Z.ai says work with safety groups in China has resulted in 2,436 vulnerability findings throughout 269 initiatives after professional evaluate, screening and deduplication. Its disclosure ledger lists 1,097 as essential or excessive severity, with 53 publicly disclosed and a pair of,383 nonetheless beneath embargo at the time of the launch.

That creates a pressure more and more dealing with frontier mannequin suppliers: the similar long-horizon agent capabilities that make fashions extra helpful for software program engineering may also make them extra succesful safety researchers — and doubtlessly extra succesful offensive operators.

GLM-5.3 additionally requires builders to change how they name the mannequin

Builders migrating present GLM purposes ought to concentrate to a breaking API habits.

GLM-5.3 helps three reasoning-effort ranges — low, excessive and max — with max the default and Z.ai’s really helpful setting for coding. However not like earlier releases, considering can’t be disabled.

Functions at the moment sending considering.sort: "disabled" should change the worth to enabled and specify a reasoning effort before switching the mannequin identifier to GLM-5.3. In any other case, Z.ai says the request will fail.

That makes GLM-5.3 an precise migration relatively than merely a model-name substitution for some manufacturing purposes.

From GLM-4.5 to GLM-5.3: Z.ai’s fast push into agentic engineering

GLM-5.3 is the newest step in a fast shift by Z.ai — previously generally known as Zhipu AI — towards coding brokers and long-running autonomous engineering workloads.

GLM-4.5, launched in July 2025, established a lot of that path. The 355-billion-parameter mixture-of-experts mannequin was designed to mix reasoning, coding and agent capabilities, whereas the smaller GLM-4.5-Air supplied 106 billion whole parameters. Z.ai launched the fashions with open weights and emphasised integration with agent frameworks.

GLM-4.6 adopted in September, increasing context from 128,000 to 200,000 tokens and focusing on coding, software use and agent workflows in environments together with Claude Code, Cline, Roo Code and Kilo Code. Z.ai additionally started putting better emphasis on token effectivity in real-world coding evaluations relatively than benchmark efficiency alone.

The bigger architectural bounce got here with GLM-5 in February 2026. Z.ai scaled the mannequin from GLM-4.5’s 355 billion parameters to 744 billion, with 40 billion energetic parameters, and elevated pretraining knowledge to 28.5 trillion tokens. It additionally launched its “slime” asynchronous reinforcement-learning infrastructure and explicitly repositioned the GLM household round “agentic engineering” and long-horizon duties.

By June, GLM-5.2 had turned that technique right into a extra direct enterprise proposition. The 753-billion-parameter mannequin arrived with a secure 1-million-token context window, open weights beneath an MIT license and help throughout greater than 20 coding environments. It additionally launched IndexShare, which reuses an indexer throughout sparse-attention layers to cut back the computational burden of very lengthy contexts.

GLM-5.2 was priced at $1.40 per million API enter tokens and $4.40 per million output tokens, with cached enter priced considerably decrease, positioning Z.ai as each a technical and pricing competitor to proprietary frontier labs.

Z.ai’s ambitions have been increasing exterior mannequin improvement as nicely. Reuters reported final month that Zhipu AI raised roughly HK$31.4 billion, or about $4 billion, by a Hong Kong share sale, with proceeds supposed for areas together with analysis and improvement, computing infrastructure, expertise and enterprise enlargement.

Taken collectively, the releases present a constant development: GLM-4.5 unified reasoning, coding and brokers; GLM-5 considerably scaled the basis mannequin; GLM-5.2 attacked long-context and long-horizon engineering; and GLM-5.3 is now making an attempt to extract considerably extra functionality from that very same basis by post-training.

Pricing, ZCode and availability

GLM-5.3 is obtainable now by Z.ai’s GLM Coding Plan and ZCode.

ZCode is the firm’s personal coding-agent setting and helps long-running “Objective” duties that plan, implement, take a look at and verify work. It additionally presents distant management of working duties and is obtainable on macOS, Home windows and Linux.

Particular person GLM Coding Plans at the moment begin at a listed promotional value of $12.60 monthly for Lite with 10,000 credit per week. Professional is listed at $56 monthly with six instances Lite utilization, whereas Max prices $117.60 monthly with 14 instances Lite utilization. Crew Normal and Premium seats are listed at $88 and $188 per consumer monthly, respectively.

Z.ai has additionally moved the Coding Plan to a points-based quota system that individually accounts for enter, cached-input and output tokens. Calls exterior the firm’s weekday peak interval devour 50% of the regular factors.

The corporate has not but supplied basic GLM-5.3 API pricing in the equipped launch supplies, making whole manufacturing API value troublesome to examine immediately with GLM-5.2 or competing frontier fashions till staged API entry arrives.

That staged launch might in the end be the most necessary a part of GLM-5.3.

Z.ai spent the previous 12 months pushing an open-model technique centered on permissive weights, low-cost inference and compatibility with present coding-agent ecosystems. GLM-5.3 demonstrates what occurs when that technique succeeds maybe too nicely in a single delicate area: higher autonomous engineering additionally means higher autonomous safety analysis.

The consequence is a mannequin that advances Z.ai’s coding ambitions whereas forcing the firm to confront the similar capability-versus-access tradeoff dealing with the largest closed frontier labs.

For enterprise builders, GLM-5.3 is subsequently price watching for 2 causes. Its coding outcomes present one other indication that more and more succesful brokers can emerge from higher post-training and environments with out constantly rebuilding the underlying basis mannequin. Its cybersecurity outcomes present why deciding how these brokers are distributed might grow to be simply as necessary as deciding how they are skilled.




Disclaimer: This article is sourced from external platforms. OverBeta has not independently verified the information. Readers are advised to verify details before relying on them.

0
Show Comments (0) Hide Comments (0)
0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments

Stay Updated!

Subscribe to get the latest blog posts, news, and updates delivered straight to your inbox.