top of page

Z.ai Tests the Anthropic Google Cybersecurity Order With GLM-5.2

Z.ai has put the anthropic google cybersecurity relationship under pressure by releasing GLM-5.2, an open-weight model with unexpectedly strong security results.

The Chinese company says its model approaches Anthropic's restricted Mythos 5 on some cyber-defense tests. Independent findings support a narrower conclusion. GLM-5.2 competes with leading closed models on selected vulnerability tasks, but it does not consistently match Mythos across broader evaluations.

That distinction matters more than the headline comparison. Google participates in Anthropic's Project Glasswing, which gives selected defenders access to Mythos-class capabilities. GLM-5.2 follows a different route. Its downloadable weights let organizations run and modify the model without a provider controlling every request.

The contest is therefore bigger than Z.ai against Anthropic. It pits restricted access and provider-managed safeguards against a model that can spread across private infrastructure. The central question is no longer whether open models will reach advanced cyber capabilities. It is how defenders respond as the gap shrinks.

GLM-5.2 Turned a Model Release Into a Security Test

Z.ai changed the debate by making credible cyber capability available through downloadable model weights.

Z.ai, also known as Zhipu AI, introduced GLM-5.2 to coding-plan users on June 13, 2026. It released the weights and technical materials three days later. The company positioned the model for coding, software engineering, and long-running agent tasks.

Open-weight means an organization can download the trained parameters and operate the model on infrastructure it controls. That arrangement differs from a hosted service, where the provider can monitor traffic, change safeguards, suspend users, or withdraw access.

The release quickly attracted attention from security researchers. Semgrep tested GLM-5.2 on insecure direct object reference detection, commonly called IDOR detection. These flaws let users access data or actions belonging to someone else.

In the IDOR evaluation, GLM-5.2 recorded a 39% F1 score. F1 combines precision and recall into one measure. Claude Code recorded 32% under the same basic prompting setup.

Semgrep's specialized multimodal pipeline remained ahead, scoring between 53% and 61%. That result supplies an important qualification. A strong model does not automatically beat a security system built around targeted analysis, repository mapping, and structured verification.

The experiment also compared different forms of assistance. GLM-5.2 received a basic harness, which is software that supplies context and manages interactions with a model. Semgrep's internal pipeline received purpose-built guidance that identified endpoints and focused the search.

This means the result does not establish broad superiority over Claude. It shows that GLM-5.2 can perform well on a reasoning-heavy vulnerability task without extensive scaffolding. That is still a meaningful change for security teams considering private deployment.

The supplied Reuters headline frames the model as nearing Mythos 5 in cyber-defense testing. Available evidence supports “nearing” only when the tested task, model configuration, and evaluation harness are clearly identified.

A benchmark result measures performance inside a defined environment. It does not automatically predict how a model handles an unfamiliar repository, a defended network, or an incomplete incident report.

GLM-5.2's importance comes from the combination of access and capability. A modestly weaker model can matter more operationally when thousands of teams can download, customize, and run it without waiting for approval.

That is why this release created immediate tension. Anthropic treats its strongest cyber functionality as a controlled capability. Z.ai has placed a competitive level of capability inside a model that can travel beyond its original developer.

Why the Anthropic Google Security Model Faces Pressure

The anthropic google approach depends on managed access, while GLM-5.2 reduces the provider's role after distribution.

Anthropic introduced Mythos Preview through Project Glasswing in April 2026. The initiative brought Anthropic together with the United States government and major technology, financial, and security organizations.

Google joined alongside Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Microsoft, Nvidia, Palo Alto Networks, the Linux Foundation, and JPMorganChase. Their stated objective was to find vulnerabilities in important software before attackers could exploit them.

Mythos 5 later became the more capable successor available to approved Glasswing participants. Anthropic describes it as the same underlying model as Fable 5, but with selected cybersecurity safeguards removed.

The Mythos release explains the division. Fable 5 serves general users with classifiers and fallbacks. Mythos 5 gives vetted defenders broader access to cyber capabilities that Anthropic considers unusually sensitive.

That structure assumes a provider remains between the model and most users. Anthropic can decide who receives unrestricted functionality. It can also monitor suspicious activity and adjust the systems surrounding the model.

Google's involvement does not make Mythos a Google model. It makes Google part of the defensive coalition using Anthropic's restricted system. The distinction is essential when interpreting the primary keyword, anthropic google, and the competitive pressure surrounding it.

GLM-5.2 challenges the coalition's operational premise. Once model weights are downloaded, Z.ai cannot reliably impose identical monitoring rules across every deployment. An operator can modify the surrounding software, fine-tune behavior, or remove refusal mechanisms.

This difference does not make open weights inherently malicious. Private operation can help defenders protect source code, regulated information, and confidential incident data. It also lets researchers reproduce findings without depending on a vendor's service.

However, the same control benefits users with offensive objectives. A malicious operator can run repeated experiments without triggering provider-side rate limits, account reviews, or centralized abuse detection.

That tradeoff puts pressure on more than Anthropic. Google and other Glasswing partners must demonstrate that controlled access produces a durable defensive advantage. They need to turn privileged capability into faster discovery, coordinated disclosure, and verified patches.

If an unrestricted competitor remains close enough, access itself becomes part of performance. A slightly stronger restricted model does not automatically provide more aggregate defense than a weaker model deployed across many internal security teams.

The pressure is both immediate and long term. In the short term, organizations must decide whether GLM-5.2 is reliable enough for security workflows. Over time, frontier developers must reconsider which safeguards still work when similar capabilities appear in downloadable systems.

This is the first major reversal in the story. Cybersecurity leadership was supposed to rest partly on limiting access to the most capable models. GLM-5.2 suggests that comparable tools can emerge outside that access model before the policy framework stabilizes.

Open Weights Change the Cybersecurity Equation

The decisive difference is not one benchmark score but who can operate, adapt, and inspect the model.

Cybersecurity work spans several distinct activities. A model might review code, reproduce a known vulnerability, develop an exploit, analyze malware, or navigate a simulated network. Success in one category does not guarantee success in the others.

Anthropic's own testing illustrates the upper end of the capability range. Mythos Preview was evaluated against patched software and thousands of open-source fuzzing targets.

According to Anthropic's cybersecurity assessment, Mythos Preview produced working Firefox exploits 181 times in a repeated experiment. An earlier Claude model succeeded twice across several hundred attempts.

The model also produced 595 lower-tier crashes in Anthropic's OSS-Fuzz testing. It reached complete control-flow hijacking, the most severe category, on ten fully patched targets.

These are company-run evaluations, so they should not be treated as universal measurements. They nevertheless show why Anthropic created a restricted access program. The relevant capability goes beyond identifying suspicious code patterns.

GLM-5.2 has not matched every one of these results in directly comparable public tests. Its current case rests on a collection of narrower findings rather than one comprehensive Mythos replication.

The United Kingdom's AI Security Institute supplies a useful counterweight to the strongest parity claims. Its researchers tested GLM-5.2 on narrow cyber tasks and simulated multi-step attacks.

The institute found that GLM-5.2 performed similarly to closed models released four months earlier on its narrow tasks. On longer cyber ranges, it resembled a model released nearly seven months earlier.

Its open-weight analysis therefore places GLM-5.2 between four and seven months behind the closed cyber frontier. That is not parity with Mythos 5.

Yet the finding is hardly reassuring for defenders relying on a lasting technical lead. The institute's previous internal evaluations placed open models six to ten months behind. The measured delay is getting shorter.

The long-horizon result deserves particular attention. GLM-5.2 initially followed a trajectory similar to a newer closed model, then stalled during the simulated attack chain.

This pattern suggests two separate capability layers. The model can solve individual technical steps, but it has more difficulty maintaining plans and adapting across an extended operation.

That limitation matters in real intrusions. Attackers must handle credentials, network changes, failed assumptions, defensive alerts, and incomplete access. A strong code-analysis score captures only part of that process.

Defenders should still avoid complacency. Open weights let outside teams improve the harness, add memory, provide specialized tools, and fine-tune the model with domain-specific examples.

A benchmark tests one packaged system at one moment. Open deployment creates a development platform. Improvements can come from organizations that never coordinate with the original model maker.

This is where GLM-5.2 changes the mechanism of competition. Anthropic can improve Mythos behind a controlled interface. The broader GLM community can improve deployment methods across many independent environments.

The anthropic google coalition retains important advantages. Its members hold extensive security data, infrastructure knowledge, and disclosure relationships. They can test findings against real systems and move fixes into widely used products.

GLM-5.2 offers a different advantage: distribution. Its users can place the model beside private repositories and customize the workflow without sending sensitive code to an outside service.

Neither advantage guarantees better security. The result depends on whether organizations can validate findings, prioritize genuine risks, and patch systems faster than adversaries can exploit them.

What the Benchmark Headlines Leave Out

The claim that GLM-5.2 nears Mythos 5 is credible on selected tasks, but false as a blanket description of cyber capability.

Security benchmarks often compress several choices into one score. Those choices include the prompt, available tools, token budget, number of attempts, model settings, and success criteria.

A model can score well because the task resembles its training data. Another can underperform because its safety system blocks part of the evaluation. A specialized harness can also outperform a stronger base model by presenting better context.

The Semgrep result demonstrates this problem clearly. GLM-5.2 beat Claude Code under a simple setup. It did not beat Semgrep's guided pipeline, which combined model reasoning with structured application analysis.

The UK institute found another limitation. GLM-5.2 approached older frontier models on narrow skills but fell further behind during long sequences. That gap weakens claims that one vulnerability-detection score represents complete offensive capability.

A recent academic test adds a third perspective. CryptanalysisBench evaluates attacks against cryptographic schemes, including known weaknesses and harder designs without established practical breaks.

On the cryptanalysis benchmark, GLM-5.2 solved 65.3% of first-tier tasks. Mythos 5 solved 85.7%. Opus 4.8, Sonnet 5, and GPT-5.5 also finished between those models.

The separation increased on harder tasks. GLM-5.2 broke 24 schemes when the researchers counted success across scaled variants. Mythos 5 broke 61.

Mythos 5 also contributed to previously unreported findings. Researchers said it identified a problem in a published security proof and helped produce a full key-recovery attack.

Those results undermine any broad claim that GLM-5.2 has already matched Mythos 5. They also show why model comparisons require multiple task families.

The evidence supports a more precise assessment. GLM-5.2 has reached a level where it can challenge leading systems on some practical security tests. Mythos 5 still holds a substantial advantage on demanding cryptographic reasoning.

This does not make the Z.ai release unimportant. A model does not need to win every benchmark before it becomes useful to attackers or defenders.

Many real security problems involve repeated, ordinary mistakes rather than advanced cryptanalysis. Access-control bugs, unsafe input handling, leaked secrets, and weak configurations appear across large software portfolios.

A widely available model that handles these tasks competently can expand the number of repositories receiving automated review. It can also increase the volume of noisy or misleading reports.

False positives impose real costs. Security teams must reproduce each finding, judge exploitability, locate affected versions, and coordinate remediation. A model that generates plausible but incorrect reports can consume scarce engineering time.

False negatives create a different risk. Teams may trust an apparently capable scanner and reduce other review activities. A benchmark victory does not justify replacing human testing, static analysis, or established fuzzing systems.

There is also a contamination concern. Public benchmark tasks can appear in training data, directly or through explanations and code. Strong evaluations use new targets, controlled environments, and trace analysis to reduce this risk.

The best response is not to select one winner from a headline. Buyers should test models on recent internal code, preserve a hidden answer set, and measure validated findings rather than generated reports.

The anthropic google story is therefore about verification as much as capability. A defensive coalition must prove that its restricted model produces operational security gains. An open model community must prove that broad access does not merely multiply unverified output.

The Real Contest Is Restricted Access Versus Distributed Capability

The competition now centers on whether safeguards can remain effective while useful cyber capability spreads.

Anthropic argues that advanced cyber models can assist both defenders and attackers. This dual-use character explains the separation between Fable 5 and Mythos 5.

Fable uses classifiers, which are separate systems that detect sensitive requests. Some flagged tasks are blocked or transferred to another model. Mythos provides approved researchers with fewer restrictions in designated areas.

This design offers several enforcement points. Anthropic can evaluate applicants, monitor traffic, investigate repeated suspicious behavior, and update classifiers when new abuse patterns appear.

The approach also creates friction for legitimate users. Security researchers often need to discuss exploitation techniques, malware behavior, or evasion methods to understand a vulnerability. A cautious classifier can interrupt that work.

Open weights remove much of this provider-controlled friction. They also remove many centralized controls. Refusal training can sometimes be altered, and an independent operator decides whether to retain logs.

The policy conflict is not simply United States versus China. It exists inside every market where model developers, regulators, and security teams balance access against misuse.

The broader security debate has already expanded beyond Mythos. Other frontier developers have introduced or tested cyber-focused systems, while researchers disagree about the timing of dangerous capability.

Google occupies a complicated position. It supports Project Glasswing as a major software and infrastructure operator. It also develops its own models and security systems, giving it incentives beyond Anthropic's access policy.

The anthropic google relationship matters because a restricted model gains value through its deployment partners. Google can supply large codebases, experienced defenders, and routes for fixing widely deployed software.

Z.ai's approach creates another form of scale. Independent organizations can place GLM-5.2 inside private development networks. They can connect it to ticketing systems, code search, test environments, and local security tools.

For enterprises, the decision involves more than raw accuracy. Data governance matters. Some organizations cannot send source code, vulnerability reports, or incident evidence to an external model service.

A locally operated model can address that constraint. However, the organization then assumes responsibility for access controls, model updates, logging, and misuse monitoring.

Deployment therefore shifts risk rather than eliminating it. Hosted services concentrate trust in the model provider. Self-hosted systems distribute responsibility across operators with widely different security practices.

This distributed pattern complicates regulation. Governments can impose conditions on domestic providers and major cloud services. They have less leverage over model copies running across private or foreign infrastructure.

It also complicates incident response. A hosted provider can patch a system-level weakness across its service. Open-model operators must obtain and apply updates themselves.

Defenders should expect both routes to persist. Highly sensitive organizations will seek controlled frontier access when it provides a measurable advantage. Others will prefer adaptable models that remain inside their environments.

The strategic mistake would be treating either route as sufficient. Restricted frontier capability cannot protect every software project. Unrestricted distribution cannot guarantee careful use or dependable findings.

A stronger defense combines capable models with traditional security controls. Those controls include code review, fuzzing, dependency management, network segmentation, identity controls, reproducible tests, and coordinated disclosure.

AI changes the speed and scale of individual tasks. It does not remove the need to decide which systems matter, confirm results, deploy fixes, and measure whether exposure actually fell.

Three Signals Will Show Whether Z.ai Has Really Closed the Gap

The next evidence should come from independent tests, field results, and measurable defensive outcomes rather than another model announcement.

The first signal is GLM-5.2's performance on broad, contamination-resistant cyber evaluations. Researchers need to test vulnerability discovery, exploitation, reverse engineering, cryptanalysis, and sustained network operations.

A stronger result across those categories would reinforce Z.ai's comparison with Mythos 5. Continued weakness on cryptanalysis or long-horizon ranges would show that the current parity claim remains task-specific.

Evaluators should publish enough methodological detail to make comparisons meaningful. They should identify the harness, tools, token limits, attempts, safeguards, and criteria used to validate success.

The second signal is how quickly open-model operators improve the surrounding system. GLM-5.2's downloadable weights allow security companies and internal teams to build specialized agents around it.

Watch for independently reproduced gains from repository mapping, retrieval, tool use, memory, or fine-tuning. If these systems close the long-horizon gap, open weights will have amplified the model beyond its release configuration.

Failure would be informative too. If optimized deployments remain unreliable, the base model's accessibility will not translate into Mythos-level operations.

The third signal is evidence of verified defensive impact. Anthropic and its Glasswing partners need to report vulnerabilities found, severity levels, duplicate rates, remediation times, and downstream patch adoption.

Z.ai users should face the same standard. Large counts mean little when reports are duplicates, false positives, low-impact issues, or flaws that maintainers cannot reproduce.

The strongest outcome would be shorter exposure windows across important software. That requires discovery, responsible reporting, engineering fixes, release management, and user adoption. Models influence only part of that chain.

Evidence of widespread offensive adaptation would change the assessment in the opposite direction. Security teams should monitor whether threat actors incorporate open models into repeatable intrusion workflows, not merely mention them in forums.

The near-term winner will not be determined by one company claiming a benchmark lead. It will be determined by who converts model capability into validated action while controlling operational risk.

For developers, the practical response is to assume that competent automated vulnerability analysis is becoming broadly available. Review high-risk code paths, improve secret handling, and make security tests reproducible before scan volume rises.

Enterprise buyers should demand evaluations on their own software rather than accepting a generic leaderboard. They should also define who validates findings and how model access is audited.

Knowledge workers tracking the anthropic google competition should preserve source material, benchmark versions, and later corrections. Capability claims change quickly, while screenshots and isolated scores often outlive their context.

The next few months will reveal whether GLM-5.2 represents narrow benchmark pressure or a lasting shift in cyber capability. Watch the independent tests, the optimized deployments, and the verified patches.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page