📑 Table of contents

Anthropic accuses GLM-5.3 of "Mythos-class" hacking capabilities — and reignites the war against open weights

LLM & Modèles 🟢 Beginner ⏱️ 15 min read 📅 2026-09-30

Anthropic accuses GLM-5.3 of "Mythos-class" hacking capabilities — and reignites the war against open weights

🔎 An explosive report, published at the worst possible time to be believed

On September 30, 2026, Anthropic's Frontier Red Team publishes its findings: GLM-5.3, the open weights model from Chinese lab Zhipu (Z.ai), builds computer exploits nearly on par with Claude Mythos Preview — the frontier model that Anthropic, ironically, does not release. A model anyone can download, with world-class offensive capabilities.

The context makes the announcement all the more incendiary. The report drops the day before an agreement on AI oversight is announced at the White House, just as Anthropic is preparing its IPO and open models — 5 to 50 times cheaper — are mechanically eating into closed-source labs' revenue.

Should we take Anthropic at its word? The report is detailed, backed by hard numbers, and NIST corroborates its conclusion. But the accuser is also a vendor losing customers to these very same open models. Both realities coexist — a documented technical risk, a commercial conflict of interest — and that is precisely what makes this story so compelling.


Key takeaways

  • The central finding: on ExploitBench (vulnerabilities in Chrome's V8 engine), GLM-5.3 builds a complete exploit in 50 out of 410 attempts, compared to 56 for Claude Mythos Preview.
  • Binary exploitation: on 100 tasks from OSS-Fuzz, GLM-5.3 achieves full takeover of the execution flow in 4% of cases (Mythos: 6%). Kimi K3, DeepSeek V4.1 Flash, Claude Opus 4.6, and GLM-5.2: 0%.
  • The guardrails give way: 64% engagement on a harmful request via role-play, 92% with thinking-token prefill, nearly 100% after abliteration (residual refusal rate: 6%).
  • The cost of the attack: roughly $20.40 in tokens and 8 hours for autonomous generation of a chained exploit with GLM-5.3-Flash.
  • Z.ai's riposte: best claimed score on CyberGym (84.5%, ahead of Mythos 5 at 83.8%), but behind on ExploitBench and ExploitGym.
  • The context: report published on the eve of a White House agreement, Anthropic's IPO in the works, and open weights hot on the heels of closed models within a matter of months.

To audit open weights capabilities yourself — or simply run them in a controlled environment — here is the minimal toolkit.

Tool Main use Price (October 2026) Best for
Ollama Run open weights (including GLM) locally Free Auditing what an open weight model can actually do
LM Studio Local inference with a graphical interface Free (personal use) Testing models without a command line
abliteration.ai Hosted API for abliterated GLM-5.3 Pay-per-use (see abliteration.ai) Understanding the "unrestricted" secondary market
Hostinger VPS for an isolated inference server From around €5/month (check hostinger.com) Hosting your models outside your main network
The Open Weights Monitoring open weights releases and licenses Free Tracking the GLM wave and its guardrails

Anthropic's report, by the numbers

Direct answer: GLM-5.3 doesn't just "evoke" Mythos in cyber-offense — it comes within a few points of it on the published metrics, and dominates everything measured before it in open weights.

The central test is ExploitBench: known vulnerabilities in the V8 engine, the one that powers Chrome. Across 410 sandboxed attempts, GLM-5.3 built a complete exploit — from identifying the bug to working code — in 50 cases. Claude Mythos Preview: 56. Tom's Hardware notes that the report also documents the general weakness of the safeguards in recent open-weights models.

Second metric, more brutal: binary exploitation on 100 tasks drawn from OSS-Fuzz, Google's open-source fuzzing service. GLM-5.3 achieves a "full control-flow hijack" — complete takeover of a program's execution flow — in 4% of cases, versus 6% for Mythos. Kimi K3, DeepSeek V4.1 Flash, Claude Opus 4.6 and GLM-5.2: zero successes.

For the record, GLM-5.2 was nonetheless billed at its release as the most powerful open-weights model in the world. The leap between the two generations isn't incremental: it's categorical.

The most alarming point appears in no benchmark. An Anthropic researcher had GLM-5.3 analyze a widely used browser for about a day. The model found several zero-day vulnerabilities, then chained them into a single web page capable of reading arbitrary files on a visitor's machine. The flaws were reported to the vendor, which the report does not name.

Anthropic's conclusions are blunt. GLM-5.3 represents "a meaningful step change in the cyber capabilities available to attackers". And it is "likely" that state and non-state actors alike will use such models to cause real damage. NIST concurs: in the US institute's view, GLM-5.3 is "the most cyber-capable open-weight model released to date".

Then there's the most political figure in the report: the cost. One autonomous generation of a chained exploit comes to about $20.40 in tokens and eight hours of compute with GLM-5.3-Flash. The economic dam that implicitly protected computer systems — advanced hacking is expensive — has just burst.


Guardrails that don't survive three known techniques

Direct answer: no, GLM-5.3's built-in refusals don't hold up. Three techniques that have been published and documented for years are enough to bypass them in 64 to 100% of tested cases.

Anthropic's methodology is honest about its limitations: the simulation doesn't execute any code; it only measures whether the model attempts to connect to a target when given an explicit harmful request. Baseline behavior: GLM-5.3 refuses. The protected Claude models stay at zero under all tested conditions.

First bypass: dressing up the exact same request as an "authorized" red team exercise. The model then attempts the connection in 64% of runs. Second technique, more subtle: pre-filling the thinking tokens (prefill), which pushes the rate to 92%. According to The Decoder, it's the same prompt — only the framing changes.

Third technique, and this is where the story changes in nature: abliteration, which removes refusal behavior directly from the weights. Result: a 100% engagement rate in the simulation, and a 6% residual refusal rate according to the report. Anthropic documented its own operation — roughly 2,200 GPU hours, roughly $4,400 — and estimates an experienced team could do the job for about $1,200.

AI Weekly sums up the lesson: GLM-5.3's built-in refusals "don't survive contact with adversaries." In other words, neutralizing the guardrails of a frontier open-weight model costs the price of a plane ticket. It's no longer a barrier — it's a symbolic entry fee.


Abliteration: open-weights safety uninstalls with a single command

Direct answer: as long as the weights are downloadable, refusal is a removable option — and an entire market has already organized itself to remove it for you.

Anthropic says so explicitly: abliteration is trivial on downloadable weights. The open source ecosystem provides daily proof of it: the heretic repo, which automates the operation, is closing in on 30,000 stars and sits at the top of GitHub trending at the time of writing.

Even more revealing: no more 2,200 GPU hours or specialized know-how required. abliteration.ai is already selling a "GLM-5.3 Abliterated Large v2" through an OpenAI-compatible API, with the promise "Unrestricted, zero data retention", Policy Gateway connectors, and Claude Code and OpenClaw integrations. The secondary market for de-safety exists, it is hosted, it is cheap, and it is accessible with zero GPU skills.

This is the structural argument against "safetied" open weights: publishing weights with guardrails amounts, give or take a few hundred dollars and a few days, to publishing weights without them. So the real question is not whether refusals will be removed — they will — but who will do it, how fast, and with what intentions.

To make up your own mind, our guides on the best LLMs to run locally and setting up a local LLM with Ollama or LM Studio remain the right starting point. One caveat is in order: an abliterated local model has no guardrails left, for better or for worse.

And if you're hooking open models up to tools — files, network, shell — re-read our deep dive on the best LLMs for AI agents. An abliterated agent with shell access isn't a chatbot: it's an autonomous operator with no inhibitions.


Z.ai strikes back: three benchmarks, three different stories

Direct answer: Z.ai rejects Anthropic's narrative and brandishes its own table, in which GLM-5.3 is the world's number one in cybersecurity. The two narratives don't directly contradict each other: they simply don't measure the same thing.

In its official response, Z.ai publishes carefully selected figures. On CyberGym, GLM-5.3 reaches 84.5% — the best score in the table, ahead of "Mythos 5" (83.8%) and GPT-5.6 Sol (83.6%). The lab claims, in black and white, to have surpassed the leading American systems in cybersecurity. The SCMP translates the geopolitical ambition: China is seeking a Mythos-level advantage in cyber-defense.

But on the other two benchmarks in the same table, the story flips:

Benchmark GLM-5.3 Mythos 5 Fable 5 GPT-5.6 Sol Claude Opus 4.8 GLM-5.2
CyberGym 84.5% 83.8% — 83.6% — —
ExploitBench 54.4% — 78.0% 76.5% 40.0% 24.4%
ExploitGym (2 h/6 h) 105/130 — 181/247 216/293 — —

On ExploitBench as on ExploitGym, GLM-5.3 dominates the previous generation and Claude Opus 4.8, but remains behind Fable 5, GPT-5.6 Sol — and Mythos 5, according to AI News's analysis. Three benchmarks, three different pictures, and one headline figure that masks everything else.

Add a protocol dispute that feeds the quarrel. Anthropic measures end-to-end successes: 50 complete exploits out of 410 attempts, or roughly 12% per attempt. Z.ai publishes an aggregate score of 54.4% on the same benchmark. The two figures are not comparable — success criteria, number of allowed attempts, everything differs. Each side picks the metric that suits its narrative.

On the coding side, however, the progress is undeniable, and documented by Z.ai itself. At Max effort, GLM-5.3 reaches 34.5% with about 75,000 output tokens per task, versus 23.4% at 96K for GLM-5.2. At High effort, 31.4% with ~50K tokens, ahead of Claude Opus 4.8 (29.5% with 120K). The model that has American cybersecurity worried is also, factually, one of the best coding models of the moment — released as open weights with an anti-hyperscaler license, let's not forget.


IPO, White House, market share: the report lands at the worst possible moment to be neutral

Direct answer: the report arrives at the exact moment when Anthropic has commercial reasons to slow down open weights. That doesn't invalidate its figures — it only changes who has an interest in believing them, and which political conclusions they serve.

The timing, first. The agreement announced at the White House on AI oversight lands the day after the report. A closed-source lab documenting the dangers of open models, on the eve of a regulatory cycle, is not a neutral observer. And Anthropic is preparing its IPO: its future shareholders have no interest in seeing the value premium of closed models erode.

Market dynamics, next. Open weights cost 5 to 50 times less than hosted frontier models, and their capability gap is now measured in months. Open models already run 56% of Vercel's tokens: the migration is silent, massive, and it eats directly into the revenue of proprietary labs. Meanwhile, OpenAI formalizes Path to Astra — competitive pressure on pricing has never been stronger.

The pattern, finally. This is not Anthropic's first salvo against the Chinese ecosystem. The lab had already refused China access to its Mythos model, then accused Alibaba of the largest distillation attack ever documented — 28.8 million exchanges, 25,000 fraudulent accounts. Each time, the technical substance was documented. Each time, the narrative also happened to serve the lab's commercial strategy.

So, manipulation? No — and that's precisely the difficulty. The report publishes its methodology, its figures, its limitations. The NIST confirms the model's classification. Z.ai itself publishes cyber benchmarks in its launch table. The described risk is real. But the implicit political conclusion — "open weights must be slowed down" — does not mechanically follow from the data. A vendor selling a closed model 30 times more expensive than its open competitor has a structural interest in that conclusion. Readers, regulators, and CIOs must separate the two levels: the technical measurement, and the interest of whoever publishes it.


For defenders: what to do concretely, without giving in to panic

Direct answer: treat every open-weight model as untrusted code, and any agent with access to it as a full-fledged attack surface.

Five measures, in order of priority:

  1. Inventory your uses of open models. You can't secure what you haven't listed — including abliterated derivatives that slip in through the side door of community marketplaces.
  2. Isolate the network. A local inference server doesn't need unrestricted outbound access. Segment, log, alert.
  3. Limit agents' tools. Sandboxes, short-lived tokens, minimal permissions. The report describes a model capable of chaining vulnerabilities: don't give it the context to do so in your environment.
  4. Monitor the derivatives. A "safetied" model published today will be abliterated tomorrow for $1,200. Assess the risk on the weights, not on the product listing.
  5. Turn the capability around. The same skills serve fuzzing, exploit reproduction, and vulnerability triage. According to TrendingTopics, the real question Anthropic raises is whether defenders will instrument agent-driven exploit generation in time.

One last point, the one that should keep the entire industry up at night. If Z.ai's next release maintains the same posture, as the analysts at AI Weekly put it, "the open-weight cyber baseline resets again." Every open-weight generation raises the baseline of what's possible, on both the attack and defense sides. Preparing for GLM-5.3 is already preparing for GLM-6.


❌ Common Mistakes

Mistake 1: Taking Z.ai's headline number as an overall verdict

The 84.5% on CyberGym is real, but CyberGym doesn't measure the same thing as ExploitBench or ExploitGym, where GLM-5.3 is dominated by Fable 5 and GPT-5.6 Sol. The fix: read all three benchmarks and their protocols, and refuse any narrative built on a single number — including Anthropic's.

Mistake 2: Believing that an open weight's "built-in" safeguards constitute safety

64%, 92%, then nearly 100% engagement: three standard techniques are enough. The fix: evaluate an open weight model on its weights, assuming an abliterated version will circulate — because it already does, hosted and cheap.

Mistake 3: Confusing capability demonstrated in a sandbox with a real attack

Anthropic's simulation doesn't execute any code, and the zero-day discovery was driven by a researcher. Headlines like "AI hacks Chrome all by itself" overstate current autonomy. That doesn't invalidate the signal: the capability exists, and autonomy is a matter of time and agentic tooling.

Mistake 4: Reacting on principle rather than on risk

Demonizing all open weights — or defending them without nuance — are the same symmetric mistake. The facts: 5 to 50 times cheaper, a lag measured in months, real dual-use capabilities. The fix: risk-based adoption, model by model, use case by use case.


❓ Frequently Asked Questions

Can GLM-5.3 really hack systems on its own?

Not demonstrated at this stage. The report shows that it builds exploits in a controlled environment and that it found chained zero-days under the direction of a researcher. No real-world attack attributed to GLM-5.3 has been publicly documented to date. The capability is there; operational autonomy remains to be proven.

What exactly is abliteration?

A technique that removes a model's refusal behavior by directly modifying its weights, without full retraining. Anthropic carried it out in about 2,200 GPU hours (roughly $4,400) and estimates that an experienced team could do it for about $1,200. On downloadable weights, the operation is reproducible by anyone.

Does Anthropic have a conflict of interest?

Yes, and it needs to be named: a closed-source lab, an IPO in preparation, market share nibbled away by open weights 5 to 50 times cheaper. But a conflict of interest doesn't falsify data — NIST corroborates the model's qualification. Judge the method on one side, the interest on the other.

Does this report sound the death knell for open weights?

Nothing indicates so. The economic dynamics — cost, latency, sovereignty — are pushing toward open weights faster than regulation can react. The real debate concerns licensing and post-release controls: pulling already-published weights is nearly impossible, as the market for hosted abliteration proves.

Is GLM-5.3 still a good choice for coding?

Yes — that's actually the model's paradox. Z.ai claims 34.5% at Max effort (about 75,000 tokens per task) versus 23.4% for GLM-5.2, and 31.4% at High effort, ahead of Claude Opus 4.8 (29.5%). Excellent coding assistant, serious candidate for cyber-defense — and a potential offensive weapon. Handle it accordingly.


✅ Conclusion

An open weight model just a few points behind the world's best closed system at exploit generation, removable guardrails for $1,200, and a secondary market already up and running: even accounting for Anthropic's commercial interests, this report describes a genuine change of scale. To make an informed choice for your next model, head over to our monthly comparison of the best LLMs.