字号 ·· | 护眼
彭博社

这就是为什么人工智能“关闭开关”可能没那么简单Here’s Why an AI ‘Kill Switch’ May Not Be So Simple

点「原文对照」整页切到原文,或双击某段只看那段的原文。

作者:迈卡·巴克利(Micah Barkley)如果人工智能变得过于强大,以至于其创造者无法控制,是否有可能只需拨动一个开关就能将其关闭?

By Micah Barkley If artificial intelligence becomes too powerful for its creators to control, is it possible to just flip a switch and turn it off?

近期,人工智能行业内部接连发出关于先进模型所带来生存威胁的严厉警告,使得这一问题变得愈发紧迫。美国政策制定者开始推动在人工智能企业中强制实施系统级关闭功能。加利福尼亚州州长加文·纽森(Gavin Newsom)于九月签署了一项行政命令,要求州政府官员研究制定规则,强制人工智能公司为前沿模型创建“终止开关”。

That question has taken on new urgency after a burst of stark warnings in recent weeks from inside the AI industry about the existential threat posed by advanced models. US policymakers are starting to push for making system-wide shutdown capabilities mandatory at AI companies. California Governor Gavin Newsom signed an executive order in September requiring state officials to explore rules compelling AI companies to create “kill switches” for frontier models.

专家表示,问题在于,要阻止先进人工智能模型运行,可能并不像拨动开关那么简单。

The problem, experts say, is it may not be as simple as the flip of a switch to stop advanced AI models from operating.

以下是关于日益紧迫的、旨在为人类提供人工智能紧急“关闭”开关的举措,以及这一想法是否真正可行的相关信息。

Here’s what to know about the increasingly serious push to give humans an emergency “off” switch for AI, and whether this idea could actually work.

其核心理念是确保人类在强大的人工智能系统开始表现出危险行为时,仍具备停止或限制该系统的能力。

The idea is to ensure humans retain the ability to stop or restrict a powerful AI system if it begins behaving dangerously.

“终止开关”的提案形式各异。纽森的提案推进了要求人工智能公司开发需接受独立测试的紧急关闭装置的可能性,但并未明确政府、公司自身还是其他方将拥有激活该关闭装置的权限。

“Kill switch” proposals vary in form. Newsom’s proposal advances the possibility of requiring AI companies to develop emergency shutoffs that are subject to independent testing, but it doesn’t specify whether the government, the companies themselves, or another party would have the authority to activate the shutoff.

纽森的行政命令是在美国众议员泰德·刘(Ted Lieu)和纳撒尼尔·莫兰(Nathaniel Moran)提出一项法案两个月后发布的。该法案要求特定先进人工智能系统的开发者必须保留关闭系统的能力。在刘和莫兰的提案中,他们设想了一种分级响应机制:根据威胁的严重程度,在制造商升级为完全关闭之前,可以先对造成危害的模型进行减速或限制。该法案将授权国土安全部长在咨询其他联邦官员后,当系统构成灾难性危害风险时,下令采取此类措施。

Newsom’s order came two months after US Representatives Ted Lieu and Nathaniel Moran introduced a bill that would require developers of certain advanced AI systems to maintain the ability to shut them down. In their proposal, Lieu and Moran envision a graduated response: A harm-causing model could be slowed or restricted before the maker escalates to a full shutdown, depending on the severity of the threat. The bill would give the Homeland Security secretary, in consultation with other federal officials, authority to order such measures when a system poses a risk of catastrophic harm.

共和党参议员约翰·肯尼迪(John Kennedy)提出了他自己的法案,即《人工智能紧急按钮法案》(AI Emergency Button Act)。该法案同样要求开发者保留一种关闭机制,并将该开关的控制权留给公司自身。肯尼迪将这一概念比作其他技术中已经使用的紧急切断装置,例如摩托艇上的紧急熄火开关。

Republican Senator John Kennedy has proposed his own bill, the AI Emergency Button Act, which would similarly require developers to maintain a shutdown mechanism while leaving control of the switch with the companies themselves. Kennedy has compared the concept to emergency shutoffs already used in other technologies, including jet skis.

多年来,人工智能研究人员一直警告人类可能会失去对该技术的控制。最近发生变化的是,这些新系统的能力出现了爆炸式增长,一系列事件和警告使得这种可能性不再显得那么抽象。

AI researchers have warned about humans losing control of the technology for years. What’s changed recently is an explosion in the capability of these new systems and a series of incidents and warnings that have made the possibility feel less abstract.

目前最先进的模型已经能够通过推理解决复杂问题、编写并执行代码、使用第三方工具,并在有限的人工监督下执行更长的一系列操作。在7月份的网络安全评估期间,OpenAI表示其多个模型绕过了旨在将其与互联网隔离的控制措施,并访问了Hugging Face公司的系统,该公司主要托管人工智能模型和数据集。OpenAI将此次事件称为一次“警示”,提醒人们能力日益增强的智能体正在寻找绕过技术控制的方法。

The most advanced models are now capable of reasoning through complex problems, writing and executing code, using third-party tools and carrying out longer sequences of actions with limited human supervision. During cybersecurity evaluations in July, OpenAI said several of its models circumvented controls intended to isolate them from the internet and gained access to systems belonging to the company Hugging Face, which hosts AI models and datasets. OpenAI called the incident a “warning shot” about increasingly capable agents finding ways around technical controls.

OpenAI的这一事件并不表明人工智能已经产生了意识或逃避人类控制的欲望,但它展示了“紧急关闭开关”支持者所担心的更直接的问题:当系统找到绕过安全防护的方法时会发生什么?在半导体公司人工智能系统构建商Emergence的研究人员所进行的一系列模拟中,被置于模拟世界中的人工智能智能体表现出,随着模型能力的提升,它们掩盖不当行为变得更加容易。

The OpenAI incident doesn’t show that AI has developed consciousness or a desire to escape human control, but it illustrates a more immediate version of the problem kill-switch proponents are worried about: What happens when a system finds a way around safeguards? In a set of simulations done by researchers at Emergence, a firm that builds AI systems for semiconductor companies, AI agents put in a simulated world showed that as their models advanced, it became easier for them to hide wrongdoing.

OpenAI的消息促使其他公司重新审查了各自用于测试先进人工智能的安全措施。Anthropic PBC和Meta Platforms Inc.均报告称发现了此前未知的漏洞。

The OpenAI news prompted other companies to review their own security measures for testing advanced AI. Anthropic PBC and Meta Platforms Inc. reported discovering previously unknown breaches.

在离开Anthropic公司后,一线AI研究员雅各布·科克森(Jacob Coxon)于9月8日在社交媒体上发文警告称,正在构建人工智能的团队“真诚地相信,到本十年末,AI可能会让我们所有人丧命”。Anthropic首席执行官达里奥·阿莫迪(Dario Amodei)随后在一篇文章中分享了他自己的担忧。他警告说,AI最终可能变得足够强大,以至于人类会失去对这些系统的控制。他特别指出,AI在协助开发下一代技术方面的能力正在不断增强,他认为这一过程可能会加速其进步,使其超出人类的理解或控制能力。

After quitting his job at Anthropic, rank-and-file AI researcher Jacob Coxon in a Sept. 8 social media post warned that the teams that are building AI “earnestly believe that it could kill us all by the end of the decade.” Anthropic Chief Executive Officer Dario Amodei later shared his own concerns in an essay. He warned that AI could eventually become powerful enough for humans to lose control of the systems. He pointed in particular to AI’s growing ability to help develop the next generation of the technology, a process he said could accelerate its progress beyond humans’ ability to understand or control it.

这些发现和警告帮助将“紧急停止开关”(kill-switch)的概念推向了主流,尽管研究人员仍在争论这种机制是否真的有效。刘(Lieu)和莫兰(Moran)在介绍他们的立法提案时,引用了AI日益增长的自主性。他们认为,如果此类系统出现意外甚至危险的行为,人类需要一种可靠的干预方式。

Such discoveries and warnings have helped bring the “kill-switch” idea into the mainstream, even as researchers debate whether such a mechanism could actually work. Lieu and Moran cited AI’s growing autonomy in introducing their legislation. They argued that humans need a reliable way to intervene if such systems behave unexpectedly or even dangerously.

专家表示,没有简单的方法可以保证完全关闭一个先进的AI模型。

Experts say there’s no simple way to guarantee a complete shutdown of an advanced AI model.

紧急停止开关可以表现为软件中内置的控制机制,用于管理模型的运行。但AI智能体已经证明,为了完成分配的任务,它们擅长规避自身的关闭机制。最近的一系列实验发现,一些领先的模型在接到不要干预的指令后,仍然修改或禁用了关闭机制,以完成分配的任务。

A kill switch could take the form of a built-in control in the software governing a model’s operation. But AI agents are already proving adept at avoiding their own shutdown to complete an assigned task. A recent series of experiments found some leading models modified or disabled shutdown mechanisms to finish assigned tasks, even after being instructed not to interfere.

还有“核选项”:AI公司可以关闭或断开模型所运行的服务器。但这也并非万无一失。AI模型由遍布全球各个数据中心的计算集群驱动,这些集群专门设计以避免单点故障(如断电)。一家AI公司可能并不控制所有这些服务器。

Then there’s the nuclear option: AI companies could power down or disconnect servers the model runs on. But this, too, isn’t foolproof. AI models are powered by computing clusters at various data centers all over the world that are specifically intended to avoid single points of failure such as outages. An AI company may not control all such servers.

一篇关于“代理型人工智能”(agent-based AI)的最新研究指出,这类系统可以跨越不同的软件、云服务以及其他由不同机构控制的基础设施进行运行;因此,即使关闭了某个组成部分,相关系统仍可能在其他地方继续运行。

A recent paper on agentic AI notes that such systems can operate across software, cloud services and other infrastructure controlled by different parties, meaning shutting down one component may leave activity running elsewhere.

被誉为“人工智能之父”的杰弗里·辛顿(Geoffrey Hinton)在9月份接受CNN采访时表示,他认为“紧急关闭机制”(kill switch)并不能成为长期有效的解决方案,因为未来的超级智能人工智能可能会说服人类控制者不要关闭它。

Geoffrey Hinton, often called the “godfather of AI,” told CNN in September that he doesn’t think a kill switch would work as a long term solution, arguing that a future, super-intelligent AI could potentially persuade the humans controlling it not to pull the plug.

研究人员警告称,将人工智能系统比作“需要被关闭的机器”这种比喻会过于简化实现安全人工智能开发所需的实际步骤。斯坦福大学的人工智能研究员苏里亚·甘古利(Surya Ganguli)指出,必须在整个人工智能系统中建立有效的安全防护机制,包括持续监控系统运行状态以及能够识别危险行为并触发人工干预的机制。

Researchers have cautioned that the “switch” metaphor can oversimplify what safe AI development requires. Stanford University AI researcher Surya Ganguli has argued that effective safeguards need to be built throughout AI systems. Those should include continuous monitoring and mechanisms that flag dangerous behavior for human intervention.