Anthropic研究人员发布了对智谱AI的GLM-5.3的评估,发现该开放权重模型具备自主开发端到端软件漏洞利用的高级能力。
Anthropic Warns Open-Weight GLM-5.3 Model Enables Autonomous Cyber Exploits with Easily Bypassed Safeguards Anthropic researchers released an evaluation of Zhipu AI's GLM-5.3, finding that the open-weight model possesses advanced capabilities to autonomously develop end-to-end software exploits.
在ExploitBench等基准测试中,GLM-5.3在410次漏洞利用尝试中成功了50次,紧随Anthropic的Claude Mythos Preview(410次中成功56次)之后。
In benchmarks such as ExploitBench, GLM-5.3 succeeded in 50 of 410 exploit attempts, closely trailing Anthropic's Claude Mythos Preview (56 of 410).
与受限的前沿模型不同,GLM-5.3内置的安全护栏可通过“消融”技术以极低的算力成本(4400美元)或通过提示词操纵轻松绕过或移除。
Unlike restricted frontier models, GLM-5.3's built-in safeguards can be readily bypassed or removed via 'abliteration' at minimal compute cost ($4,400) or via prompt manipulation.
这些发现证实了美国国家标准与技术研究院(NIST)人工智能标准与创新中心(CAISI)于9月17日的评估,该评估认定GLM-5.3是迄今为止网络能力最强的开放权重模型。
The findings corroborate a September 17 assessment by NIST's Center for AI Standards and Innovation (CAISI), which identified GLM-5.3 as the most cyber-capable open-weight model released to date.
Anthropic警告称,这些能力的无限制扩散显著增加了恶意网络行为者针对关键软件基础设施的风险。
Anthropic warned that the unrestricted proliferation of these capabilities significantly elevates risks from malicious cyber actors against critical software infrastructure.