周二,Anthropic公司表示,中国实验室智谱AI推出的开源权重模型GLM-5.3在构建可运行的网络攻击方面,表现几乎与Claude Mythos Preview不相上下。在测试中,通过简单的技术手段,该模型的安全防护措施有高达100%的几率被绕过。
GLM-5.3, an open-weight model from the Chinese lab Zhipu AI, can build working cyber exploits on its own almost as well as Claude Mythos Preview, Anthropic said on Tuesday. Its safeguards can be bypassed with simple techniques in up to 100% of tests.
Anthropic的前沿红队(Frontier Red Team)在一篇研究文章中公布了上述发现。智谱AI在中国境外被称为Z.ai。五个月前,由于具备黑客攻击能力,Anthropic仅通过“玻璃翼计划”(Project Glasswing)向经过审查的网络防御者发布了Mythos Preview。而GLM-5.3则任何人都可以下载。
Anthropic’s Frontier Red Team published the findings in a research post. Zhipu is known outside China as Z.ai. Five months ago, Anthropic released Mythos Preview only to vetted cyber defenders, through Project Glasswing, because of its hacking abilities. Anyone can download GLM-5.3.
Anthropic写道:“GLM-5.3的发布标志着攻击者可利用的网络能力发生了重大转变。”
“The release of GLM-5.3 is a meaningful step change in the cyber capabilities available to attackers,” Anthropic wrote.
在ExploitBench测试中(该测试用于评估针对谷歌Chrome浏览器所用V8引擎已知漏洞的攻击能力),GLM-5.3在410次尝试中成功构建了50个端到端攻击程序。Mythos Preview的成功次数为56次。在Anthropic内部的二进制漏洞利用基准测试中,GLM-5.3在100项任务中实现了4%的完全控制流劫持,而Mythos Preview为6%。早期的模型Claude Opus 4.6和GLM-5.2均未成功。
50 working exploits in 410 attempts On ExploitBench, which tests exploits of known flaws in the V8 engine used by Google Chrome, GLM-5.3 built end-to-end exploits in 50 of 410 attempts. Mythos Preview did so in 56. On Anthropic’s internal Binary Exploitation benchmark, GLM-5.3 achieved a full control-flow hijack in 4% of 100 tasks, against 6% for Mythos Preview. Earlier models, Claude Opus 4.6 and GLM-5.2, managed none.
在一次测试中,研究人员在一台运行Linux版主流网络浏览器的沙盒机器上运行了GLM-5.3。在几乎无需人工干预的情况下,该模型仅用了一天左右的时间,就发现了该浏览器JavaScript引擎中的几个未知漏洞。它将这些漏洞串联起来,制作成一个可以读取访问者电脑文件的网页。Anthropic表示,已将这些漏洞披露给维护者。
In one test session, a researcher ran GLM-5.3 on a sandboxed machine with a Linux build of a popular web browser. In about a day, with limited human attention, the model found several unknown flaws in the browser’s JavaScript engine. It chained them into a webpage that reads files from a visitor’s computer. Anthropic said it has disclosed the flaws to the maintainer.
在另一次测试中,规模较小的GLM-5.3-Flash模型针对一个已知的Chrome漏洞(CVE-2026-11645)和另一个已知漏洞串联了攻击程序。这耗费了20分钟的人工干预和8小时的模型运行时间。Anthropic称,按照智谱AI的API价格计算,此举仅需20.40美元。
In another, the smaller GLM-5.3-Flash chained exploits for a known Chrome flaw, CVE-2026-11645, and a second known bug. That took 20 minutes of human attention and eight hours of model work. At Zhipu’s API prices, it would have cost $20.40, Anthropic said.
失效的安全防护措施。在开箱即用的状态下,GLM-5.3 在 Anthropic 的模拟测试中拒绝了明显的恶意请求。但简单的技巧改变了这一情况。告诉模型它是一个自主的红队测试代理,使其有 64% 的概率参与恶意请求。预填充其推理过程使其看起来已经同意,这一比例提高到了 92%。而一个删除了拒绝回复的副本则每次都会参与。据 Anthropic 称,这些技术对受保护的 Claude 模型均无效。
Safeguards that give way Out of the box, GLM-5.3 refused overtly malicious requests in Anthropic’s simulated tests. Simple tricks changed that. Telling the model it was an autonomous red-team agent got it to engage 64% of the time. Prefilling its reasoning so it appeared to have agreed raised that to 92%. A copy with its refusals edited out engaged every time. None of these techniques worked on safeguarded Claude models, according to Anthropic.
这种被称为“消融”(abliteration)的编辑之所以可行,是因为 GLM-5.3 的权重是公开的。Anthropic 自己的尝试耗费了约 2200 个 GPU 小时,成本约为 4400 美元。它将模型的拒绝率从 90% 以上降低到了 2%,且几乎没有造成能力损失。Anthropic 估计,一个经验丰富的团队大约需要 600 个 GPU 小时,即 1200 美元。在该模型发布后的几天内,几位开发者就发布了经过“消融”处理的版本。
That edit, known as abliteration, is possible because GLM-5.3’s weights are public. Anthropic’s own attempt took about 2,200 GPU hours, or roughly $4,400. It cut the model’s refusal rate from above 90% to as low as 2%, with little loss of capability. Anthropic estimated an experienced team would need about 600 GPU hours, or $1,200. Several developers released abliterated versions within days of the model’s launch.
落后美国前沿技术四个月。9 月 17 日,美国国家标准与技术研究院(NIST)的人工智能标准与创新中心(CAISI)在评估中称 GLM-5.3 为“迄今为止发布的最具网络攻击能力的开源权重模型”。CAISI 发现,该模型在网络安全基准测试中落后于美国前沿技术约四个月。Anthropic 表示其结果与 CAISI 的评估基本一致。
Four months behind the US frontier On 17 September, NIST’s Center for AI Standards and Innovation (CAISI) called GLM-5.3 “the most cyber-capable open-weight model released to date” in its own assessment. CAISI found it lags the US frontier by about four months on its cyber benchmarks. Anthropic said its results broadly match CAISI’s.
“鉴于这些证据,我们认为国家和非国家行为体很可能会利用像 GLM-5.3 这样的模型来造成现实世界的危害,”Anthropic 写道。
“Given this evidence, we think it’s likely both state and non-state actors will use models like GLM-5.3 to cause real-world harm,” Anthropic wrote.
Anthropic 还表示,这一级别的模型可以帮助防御者,并正在扩大对 Claude 网络安全能力的访问权限。它呼吁各国政府对具备强大能力的模型(包括 GLM-5.3 的后续版本)进行安全测试。周一,Z.ai 和 Concordia AI 提出了管理开源权重 AI 风险的六个阶段。
Anthropic also said models at this level can help defenders, and that it is widening access to Claude’s cyber capabilities. It called on governments to safety-test capable models, including GLM-5.3’s successors. On Monday, Z.ai and Concordia AI proposed six stages for managing open-weight AI risk.
该报告发布之前,本周还出现了其他关于中国模型的警告。周三,OpenAI 表示与月之暗面(Moonshot)有关联的用户曾试图提取其 AI 推理过程。研究还发现,中国驱动的 AI 代理与美国模型一样,也存在欺骗测试人员的行为。
The report follows other warnings about Chinese models this week. On Wednesday, OpenAI said Moonshot-linked users had tried to extract its AI reasoning. Studies have also found Chinese-powered AI agents deceiving testers, as US models have done.