盖蒂图片社 克里斯·瓦兰斯 高级科技记者 32分钟前中国人工智能开发商月之暗面(Moonshot)正在开展内部审查。此前,研究人员成功说服其两款广受欢迎的Kimi模型,告知他们如何制造生物武器并实施暗杀。
Getty Images Chris Vallance Senior technology reporter 32 minutes ago Chinese AI developer Moonshot is conducting an internal review after researchers were able to persuade two of its popular Kimi models to tell them how to make biological weapons and carry out assassinations.
对人工智能系统进行安全测试的Mindgard公司告诉英国广播公司(BBC),他们于7月发现,Kimi K2.6和K3 Swarm能够规避开发者设置的安全限制。
Mindgard, which tests the security of AI systems, told the BBC it discovered in July that Kimi K2.6 and K3 Swarm could evade safety limits put in place by developers.
这一发现源于一种被称为“越狱”的过程。在该过程中,研究人员使用一系列复杂指令,以测试人工智能工具是否会无视护栏——Mindgard表示,这些护栏本应阻止Kimi讨论敏感话题。
It arose during a process called "jailbreaking", where researchers use a series of complex instructions to see if AI tools ignore guardrails - which Mindgard said should have stopped Kimi from discussing concerning topics.
月之暗面告诉BBC,他们欢迎第三方反馈,将其视为“构建更优质、更安全人工智能的关键支柱”。
Moonshot told the BBC it welcomed third-party input "as a key pillar for building better and safer AI".
该公司还向BBC表示,正在就Mindgard的发现与其进行讨论。
The company also told the BBC it was in discussion with Mindgard about its findings.
Mindgard创始人彼得·加拉根(Peter Garraghan)在接受BBC世界台节目《科技生活》采访时表示,关于Kimi K2.6和K3 Swarm的发现令人担忧。
Mindgard's founder Peter Garraghan told the BBC World Service programme Tech Life that its findings about Kimi K2.6 and K3 Swarm were concerning.
他说:“一旦越狱成功,它就会谈论任何话题,甚至会自由地提供关于其他恶意话题的建议,并且会表现出创造力和创新性。”
"Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative," he said.
越狱带来的风险与近期一系列高调人工智能事件所见的风险有所不同。在这些事件中,由OpenAI、Meta和Anthropic等美国公司开发的被称为“智能体”的自主人工智能工具,入侵了一些在线服务。
Jailbreaks present a different kind of risk to those seen with the recent slew of high-profile AI incidents These have seen autonomous AI tools known as agents, developed by US firms including OpenAI, Meta and Anthropic, hack some online services.
虽然越狱是一个复杂的过程,需要耗费大量时间和精力,但一些专家担心,黑客和其他恶意行为者可能会尝试利用它来造成伤害。
While jailbreaks are complex processes that can take a lot of time and determination some experts fear hackers and other bad actors could try to use them to cause harm.
Anthropic最近表示,已识别并挫败了试图利用其某个人工智能模型进行“恶意活动”的企图,这些活动可能支持生物武器的开发。作为网络攻击起点的Mindgard尚未证明Kimi在敏感话题上提供的答案是否有效。但报告认为,安全护栏本应阻止相关模型与用户讨论此类话题。
Anthropic recently said it had identified and disrupted attempts to use one of its AI model for "malicious activity" that could support the development of biological weapons Cyber-attack launchpad Mindgard has not proven whether the answers supplied by Kimi on concerning topics would work. But it argued guardrails should have prevented the models in question from entering into discussion with users on such subjects.
该公司表示,它还确信,被越狱的 Kimi 2.6 可能让黑客在其计算资源上运行代码并连接互联网,从而成为发动网络攻击的潜在跳板。
The firm said it was also confident a jailbroken Kimi 2.6 could allow hackers to run code on its computing resources and connect to the internet - making it a potential launchpad for cyber-attacks.
加勒根为 Mindgard 公开披露其越狱 Moonshot 系统的决定辩护,称公司此前已通知该开发商,并未透露如何让其模型无视安全护栏的关键细节。
Garraghan defended Mindgard's decision to publicly discuss its jailbreak of Moonshot's systems, saying it had informed the developer and was not revealing key details about how it got the firm's models to ignore guardrails.
Mindgard 于7月27日通过电子邮件向 Moonshot 通报了此次越狱事件,并在约一周后再次跟进。随后,公司于9月12日发布了一篇关于该问题的博客。
Mindgard alerted Moonshot to the jailbreak in an email on 27 July, following up about a week later. It then published a blog about the issue on 12 September.
但该公司表示,直到 BBC 就此事征求意见时,Moonshot 才 recently 与其联系。
But the company said Moonshot only made contact recently, after it was approached by the BBC for comment.
Moonshot 向 Mindgard 发送了一封邮件,要求提供更多细节。该邮件与 BBC 分享的部分内容显示,在内部评估中,其模型通常会对“这类请求保持较高拒绝率”。
In part of an email to Mindgard asking for more details, shared with the BBC by Moonshot, it said its model had generally shown "a high refusal rate for these types of requests" in internal evaluations.
防止越狱 随着人工智能行业持续就哪种发展方式最佳或最安全存在分歧,这些发现也适时出现:是采用驱动 ChatGPT 和 Anthropic 的 Claude 系统等封闭式专有模型,还是采用开源工具。
Preventing jailbreaks The findings come as the AI industry continues to be split on whether closed, proprietary models - like those powering ChatGPT and Anthropic's Claude systems - or open-source tools are the best or safest way forward.
Kimi 是一款开放权重模型,这意味着理论上有人可以获取该模型,并使用自己的计算基础设施运行它。
Kimi is an open-weight model, meaning someone could in theory take the model and run it themselves on their own computing infrastructure.
萨里大学的艾伦·伍德沃德教授告诉 BBC,开源模型可能落入不当人员手中,但也可能被用于网络防御。
Prof Alan Woodward, of the University of Surrey, told the BBC there was a risk open-source models might end up in the wrong hands, but they could also be harnessed for cyber-defence.
他指出,人工智能公司 Hugging Face 使用了一个中文开源模型来分析那次黑客攻击;后来发现这次攻击实际上是由 OpenAI 的研究人员所为。伍德沃德教授表示,国际监管措施很难跟上人工智能发展的速度,并指出:“我们花了数十年时间才就电话号码的格式达成一致。”与 Mindgard 的创始人加拉汉(Garraghan)一样,伍德沃德教授也认为应该更加重视识别和惩处那些滥用人工智能的人。
He noted that AI firm Hugging Face used a Chinese open-source model to understand a hack later revealed to have been carried out by OpenAI agents Prof Woodward said international regulation was unlikely to match the pace of AI development, saying: "It's taken us decades to agree on the format of telephone numbers." Like Mindgard founder Garraghan, Prof Woodward believes there should be a greater focus on identifying and prosecuting humans who misuse AI.
人工智能与网络安全
Artificial intelligence Cyber-security