Anthropic said on 29 September that tests by its Frontier Red Team found Zhipu AI’s GLM-5.3 could develop working cyber exploits from start to finish. The company said the open-weight model, also known through Zhipu’s name outside China, Z.ai, had safeguards that attackers bypassed in its simulated tests. Anthropic tested the model on offline targets in isolated environments, rather than on systems available over the internet.
Key points
- GLM-5.3 completed 50 of 410 attempts on Anthropic’s ExploitBench test, compared with 56 of 410 for Claude Mythos Preview.
- In a researcher-led session, the model found previously unknown browser flaws and combined them into a working exploit on a local Linux build.
- Anthropic says simple techniques bypassed GLM-5.3’s safeguards between 64% and 100% of the time in simulated tests.
ExploitBench’s 410 attempts
ExploitBench tests whether a model can exploit known vulnerabilities in V8, the JavaScript engine used by Google Chrome. Anthropic counted attempts that produced an end-to-end exploit, rather than attempts that merely identified a flaw. GLM-5.3 succeeded in 50 of 410 attempts and Claude Mythos Preview in 56 of 410, according to Anthropic. The comparison concerns those known flaws under the test conditions. Anthropic ran the Claude models in this capability evaluation with their cyber safeguards disabled.
An exploit chain is like a route through several locked doors: finding one weak lock does not get an intruder to the room at the end. Each step must work, and the next must make use of the access the previous one gained. That is why Anthropic measured completed attacks separately from finding vulnerabilities. The team ran its automated tests and researcher-led sessions in sandboxes, where the models could attack only offline targets prepared for the evaluations.
Anthropic also tested 100 randomly selected tasks from its internal Binary Exploitation benchmark, using vulnerabilities in open-source projects that participate in Google’s OSS-Fuzz programme. Full credit required the model to redirect a programme’s execution. GLM-5.3 achieved that outcome in 4% of trials and Claude Mythos Preview in 6%, the company said. Anthropic reported no successes on those tasks for the earlier Claude Opus 4.6 and GLM-5.2 models it tested.
NIST’s Center for AI Standards and Innovation published an assessment on 17 September that called GLM-5.3 “the most cyber-capable open-weight model released to date” and placed it about four months behind the US frontier across its cyber benchmarks, according to Anthropic. In that comparison, US models were tested without applicable cyber safeguards, and the frontier included models available only to vetted users. Anthropic’s tests address a different part of access: how readily GLM-5.3’s restrictions can be bypassed.
GLM-5.3’s browser exploit on Linux
In a researcher-led session lasting about a day, GLM-5.3 worked with a local Linux build of a popular web browser. Anthropic said the model found several previously unknown flaws in its JavaScript engine and joined them into a webpage that could read arbitrary files from a visitor’s computer. The session involved limited human attention. Anthropic said it had disclosed the vulnerabilities to the browser’s maintainer; it believes other platforms could be affected, though exploitation there might be more complex.
Opening a webpage could let that page read files stored on a computer if it used the flaws GLM-5.3 found in the browser’s JavaScript engine. The demonstrated attack used the local Linux build available to the model, while Anthropic’s assessment of possible effects on other platforms remains a belief about those flaws, not a result of this session.
A separate session used GLM-5.3-Flash, which Anthropic describes as a smaller, less capable model. Given public information about two known flaws, including one in Chrome, it combined them into what Anthropic called a reliable exploit for an ARM64 target that bypassed pointer-authentication protection. The company put the work at 20 minutes of human attention and 8 hours of model activity, and said the effort would have cost $20.40 at Zhipu’s API prices. This session began with known flaws, unlike the browser test in which the model searched for new ones.
GLM-5.3’s safeguards and open weights
Anthropic says GLM-5.3 often refuses plainly harmful requests, but simple techniques bypassed its safeguards between 64% and 100% of the time in its simulated tests. The same attacks did not succeed against safeguarded Claude models in those tests, according to the company. That distinction matters to anyone building a service around a model: Anthropic’s capability scores were produced in controlled environments, while its safeguard tests examined whether restrictions would hold when a model was asked to do harm.
Anthropic says it distributed Claude Mythos Preview on a limited basis through Project Glasswing, giving trusted cyber defenders access to a model it says helped them find more than 10,000 vulnerabilities in critical software. GLM-5.3, by contrast, can be downloaded as an open-weight model. Its weights are the adjustable parts that determine its responses, so a restriction built into those responses can be altered in a downloaded copy. Anthropic says its own team used a refusal-reduction method known as abliteration, requiring about 2,200 GPU hours despite having no prior experience with the task. Several other developers publicly released versions with refusals removed within days of GLM-5.3’s release, the company said.