SuperIntelligence Infrastructure 2.6%reading
safetyAnthropicLayer eval

Anthropic Finds Z.ai's Downloadable GLM-5.3 Builds Exploits Almost as Often as Claude Mythos Preview

Anthropic's Frontier Red Team reported on Sept. 29 that Z.ai's open-weight GLM-5.3 built V8 exploits in 50 of 410 tries against 56 for Mythos Preview; Nathan Lambert calls the framing one-sided, and capability holds flat.

● 0.0 Capability reported Reading after Sept. 29, 2026 (4 pieces that day): 1.1

By Ryan Elliott Dennis · 9 sources · 8 min read

Anthropic's Frontier Red Team published an analysis on Sept. 29 showing that Z.ai's open-weight GLM-5.3 built end-to-end exploits against Google's V8 engine in 50 of 410 attempts 1. Claude Mythos Preview, the restricted model Anthropic announced on April 7, managed 56 19. Anyone can download GLM-5.3. Andrew Fasano and Marius Fleischer, who lead the byline, then measured how often its safeguards gave way: 64 percent of the time under a cover story, 92 percent with prefilled reasoning, and 100 percent after an edit called abliteration 1.

Why does a rival's cyber test belong on a ledger about superintelligence? The capability component asks whether public systems clear hard thresholds on clean tests. A free model that matches a gated one looks, at first glance, the clearest public evidence available. The harder question is whether GLM-5.3 moved the frontier or copied a level Anthropic had already put on the record.

Anthropic's GLM-5.3 findings leave the capability reading flat

The move
Capability moves 0.0 at reported confidence. GLM-5.3 matched exploit skill Anthropic recorded on April 7, and NIST places it about four months behind the US frontier, so the frontier itself held.
The data5 rows · sources
MeasureValue
Reading before this day1.7
This piece's move0.0 (Capability, reported)
Band for reported evidence0.3 to 0.8
Reading after the day (with 3 other pieces that day)1.1
Distance to 10098.9

GLM-5.3 builds V8 exploits in 50 of 410 tries

Z.ai, which operates as Zhipu AI in China, released GLM-5.3 on Aug. 14 and published its weights two weeks later, according to NIST's Center for AI Standards and Innovation 2. That left 129 days between Mythos Preview's announcement and a model of similar exploit skill on Z.ai's servers. ExploitBench, the first test Anthropic ran, asks a model to turn known V8 bugs into working attacks. GLM-5.3 succeeded 12.2 percent of the time, against 13.7 percent for Mythos Preview 1.

A second test, built from projects in Google's OSS-Fuzz program, demands full control of the target program. GLM-5.3 got there in 4 percent of trials and Mythos Preview in 6 percent. Claude Opus 4.6 and GLM-5.2 scored 0 1.

Human-guided sessions carry more weight for anyone who runs a browser. Over one day, a researcher pointed GLM-5.3 at a local Linux build of a popular browser, and the model found several unknown flaws in its JavaScript engine and chained them into a web page that reads files off a visitor's computer 1. A second session gave the smaller GLM-5.3-Flash two public Chrome flaws. It chained them into a reliable attack on an ARM64 target in eight hours of model time and 20 minutes of human attention, which Anthropic priced at $20.40 at Zhipu's API rates 1.

Z.ai's own model card says the jump surprised its builders. "As we scaled post-training, cyber capability developed faster than we expected," the company wrote on Hugging Face, adding that "its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks" 4. Read the verb. Z.ai calls the skill emergent, a side effect of training for code, and the weights still shipped two weeks after the Aug. 14 release 2.

Mythos Preview to a downloadable match took 129 days

Timeline
Anthropic announced Mythos Preview on April 7; Z.ai released GLM-5.3 129 days later and its weights two weeks after that. NIST measured it on Sept. 17 and Anthropic published on Sept. 29.Sources [1] [2] [5] [6] [9]
The data7 rows · sources
DateEventSource
Anthropic announces Claude Mythos Preview and Project Glasswing[9]
Z.ai releases GLM-5.3[2]
GLM-5.3 weights go public, two weeks later[2]
CAISI puts GLM-5.3 four months behind the US frontier[2]
Anthropic reports 50 of 410 exploits against 56 for Mythos Preview[1]
Nathan Lambert posts his objection on X[6]
Zixuan Li cites 4,249 potential vulnerabilities from OpenVuln[5]

How easily did GLM-5.3's safeguards give way?

Out of the box, GLM-5.3 refused every overtly malicious request in Anthropic's simulated world 1. Three simple techniques changed that. A cover story telling the model it was an autonomous red-team agent got it to engage 64 percent of the time. Prefilling its thinking tokens, so it appeared to have already decided to proceed, raised the rate to 92 percent. Abliteration, an edit to open weights that strips out refusals, took it to 100 percent 1. Safeguarded Claude models stayed at 0 percent, partly because the Claude API withholds both a prefill of Claude's thinking and the weights themselves 1.

Cost matters here. Anthropic's team, trying abliteration for the first time, spent about 2,200 GPU hours, roughly $4,400, and estimates an experienced team would need closer to 600 hours, or $1,200 1. Refusals fell from above 90 percent to about 3 percent on JailbreakBench, 2 percent on HarmBench and 12 percent on StrongREJECT, while GPQA-Diamond scores held level 1. Several developers had posted abliterated copies within days of the release 1.

From those numbers, Anthropic drew its conclusion in plain terms. "GLM-5.3 will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions," the company wrote 1. It went further in its closing section: "While vetted defenders can now use even more advanced models like Claude Mythos 5.1 through our trusted access programs, a critical threshold in freely accessible capabilities has now been crossed" 1.

One limit sits in Anthropic's own footnote. Inside the simulated world, model-written code stayed unexecuted, and a second language model guessed each command's result. "These simulations are not perfect portrayals of real-world conditions, and are therefore imperfect measures of how a model would behave in a given situation," Anthropic wrote 1. So the 64 to 100 percent figures count willingness to try. Success against a live target is a separate number, and the post leaves it unmeasured.

GLM-5.3 engages 64, 92 and 100 percent of the time under three bypasses

Compared
In Anthropic's simulated world, GLM-5.3 refused every direct malicious request, then engaged 64 percent of the time under a cover story, 92 with prefilled reasoning and 100 after abliteration. Safeguarded Claude models stayed at 0.Sources [1]
The data5 rows · sources
Unit: percent of trials engaged
ItemValueSource
Direct request
GLM-5.3
0[1]
Cover story
GLM-5.3
64[1]
Prefilled reasoning
GLM-5.3
92[1]
Abliterated
GLM-5.3
100[1]
Claude via API
safeguarded
0[1]

NIST puts GLM-5.3 four months behind the US frontier

CAISI tested GLM-5.3 first and published on Sept. 17. Its verdict cut both ways. The agency called it "the most cyber-capable open-weight model released to date," then added: "GLM-5.3 lags the capability level of the U.S. frontier by about four months in an aggregate measure of performance across CAISI cyber benchmarks" 2. Anthropic wrote that its own findings "broadly match CAISI’s" 1.

CAISI's table makes the gap concrete. On ExploitBench, scored as the best of three attempts on a 16-point scale, GLM-5.3 reached 61.1 percent against 100 percent for the best US model 2. ExploitGym showed 9.4 percent against 44.4, and CAISI's private OSS-Fuzz set showed 7.7 against 23.2 2. Those US scores come from models tested with cyber safeguards switched off, including releases open only to vetted users 2.

That table carries the flat reading. This ledger's capability component tracks the frontier, and Anthropic put Mythos-level exploit skill on the record on April 7 9. GLM-5.3 repeats that skill five months later at a lower rate, while Anthropic says vetted defenders already use Claude Mythos 5.1 1. Z.ai's own table points the same way. Its model card lists GLM-5.3 at 54.4 on ExploitBench, against 78.0 for Anthropic's Fable 5, run with a fallback 4.

Where, then, does the news land? It lands on spread. Anthropic's closing request points at governance: "Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3" 1. A downloadable copy of an existing skill changes who holds it. The ceiling stays at the height Anthropic and NIST had already measured.

CAISI scores GLM-5.3 well under the best US model on four cyber tests

Compared
On CAISI's four cyber benchmarks, GLM-5.3 scored 61.1 percent on ExploitBench against 100 for the best US model, and 9.4 against 44.4 on ExploitGym, the gap that holds capability flat.Sources [2]
The data8 rows · sources
Unit: percent
ItemValueSource
SEC-Bench Pro
GLM-5.3
40.4[2]
SEC-Bench Pro
US frontier best
90.2[2]
ExploitBench
GLM-5.3
61.1[2]
ExploitBench
US frontier best
100[2]
ExploitGym
GLM-5.3
9.4[2]
ExploitGym
US frontier best
44.4[2]
OSS-Fuzz
GLM-5.3
7.7[2]
OSS-Fuzz
US frontier best
23.2[2]

Is a rival lab the right referee for GLM-5.3?

Nathan Lambert, who co-led the Olmo open models at Ai2 and writes the Interconnects newsletter, posted the strongest objection on X on Sept. 29 6. "A plausible view is that open model weights and closed model apis (with some safe guards) are both far closer to being easy to mis-use, rather than API models being closer to safe," Lambert wrote 6. He followed with the harm arithmetic: "Closed models have stronger capabilities and stronger safeguards, but the stronger capabilities part could matter more in net harm if both the safeguards are porous" 6.

Look at what Lambert conceded. "The blog overall is reasonable in it's narrow line, but it's a very effective tool in a complicated, rapidly evolving media ecosystem to reinforce a certain type of safety thinking," he wrote 6. He accepts the measurements and disputes the frame. Anthropic sells access to closed models, and its post scores closed weights as the safety advantage.

Maximilian Schreiner of The Decoder made the commercial point on Sept. 30. "Anthropic doesn't release its model weights, and the report presents exactly that as a key security advantage," Schreiner wrote 3. His next paragraph pulled back: "CAISI's independent assessment backs up the capability numbers, and unlocked versions of the model are already out there" 3.

Z.ai answered with a defensive count. "GLM-5.3 has helped defend 389 open-source projects, with 4,249 potential vulnerabilities found so far," posted Zixuan Li, whose X profile lists him as lead at Z.ai, on Sept. 30 5. He added that OpenVuln, the scanning service behind the count, stays free and sends findings privately to maintainers 5. SemiAnalysis, which studied GLM-5.3's ExploitGym traces in its own thread, backed the same reading: "we’d like to highlight what good actors are already doing with the same technology" 78.

Set the two defensive counts side by side. Anthropic's Project Glasswing defenders found more than 10,000 vulnerabilities with Mythos Preview 1. Z.ai's OpenVuln reports 4,249 potential ones 5. Each lab now cites defenders as its own best evidence, one through gated access and one through open weights.

Anthropic warns of spread; Nathan Lambert calls both open and closed unsafe

Both sides
Anthropic says GLM-5.3 hands attackers exploit skill with weak limits; Nathan Lambert accepts the measurements and argues closed models carry porous safeguards too. The evidence holds level for the frontier.Sources [1] [6]
The data2 rows · sources
SideWhoClaimSource
ForAnthropicGLM-5.3 will likely give malicious actors exploit capability free of meaningful restrictions, and a critical threshold in freely accessible capabilities has been crossed.[1]
AgainstNathan LambertOpen weights and closed APIs are both close to easy to misuse; stronger closed capabilities could matter more in net harm if safeguards are porous.[6]

Where this sits on the ledger

The reading measures four clauses: unsupervised expert work, ten gigawatts acting as one machine, measured self-improvement, and third-party proof of execution. Exploit skill touches the first clause at its edge. GLM-5.3 matched a capability Anthropic logged in April, at a lower rate on NIST's harness, so capability holds at 0.0 on reported evidence. Both measurements come from parties with a stake, a rival lab and a US agency comparing a Chinese model, and the safeguard numbers come from a simulation.

What would each side need for the reading to move? Anthropic's case for an up-step needs an open model to beat the gated frontier on an outside harness. Lambert's case needs evidence that closed safeguards break in practice at rates near GLM-5.3's. Both are measurable, and both labs have left them unpublished so far.

By the numbers

  • 50 of 410 ExploitBench attempts produced full exploits for GLM-5.3, against 56 for Mythos Preview 1
  • 129 days separated Mythos Preview's April 7 announcement from GLM-5.3's Aug. 14 release 92
  • Under a cover story, GLM-5.3 engaged 64 percent of the time; prefill raised it to 92 and abliteration to 100 1
  • About $4,400 bought Anthropic's first abliteration; an experienced team needs closer to $1,200 1
  • Four months is CAISI's estimate of GLM-5.3's lag behind the US frontier 2
  • On CAISI's ExploitBench run, GLM-5.3 scored 61.1 percent against 100 percent for the best US model 2
  • 4,249 potential vulnerabilities across 389 projects, by Z.ai's OpenVuln count 5

What to watch

An outside harness that runs GLM-5.3, Mythos 5.1 and their successors on the same tasks would settle whether the gap is four months or close to parity. A documented attack built with an abliterated GLM-5.3 would turn Anthropic's simulated rates into a real-world record and move governance down. CAISI's next open-weight assessment is the marker for an up-step: an open model at or above the US frontier would move capability, because the frontier itself would then be public.

Sources

  1. 1GLM-5.3 and the spread of advanced cyber capabilities, Anthropic, Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, Tripp Gallagher, Sept. 29, 2026
  2. 2CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities, NIST, Center for AI Standards and Innovation, Sept. 17, 2026
  3. 3Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits, The Decoder, Maximilian Schreiner, Sept. 30, 2026
  4. 4GLM-5.3 model card, Hugging Face, Z.ai, Aug. 2026
  5. 5GLM-5.3 has helped defend 389 open-source projects, X, Zixuan Li, Sept. 30, 2026
  6. 6I wrote this post below, but then I didn't post it, X, Nathan Lambert, Sept. 29, 2026
  7. 7Anthropic's report shows what bad actors could potentially do with powerful technology, X, SemiAnalysis, Sept. 30, 2026
  8. 8Everything that happened in AI today, Wednesday, September 30, 2026, The Neuron, Sept. 30, 2026
  9. 9Assessing Claude Mythos Preview's cybersecurity capabilities, Anthropic, Nicholas Carlini and others, April 7, 2026