如果行业在安全问题上不做出改变,其当前的做法只会导致更多失败。
The industry’s approach to safety will guarantee more failures unless something changes.
我意识到,接下来要告诉各位的话,已经成了一种陈词滥调:我本周辞去了在OpenAI的职务。在每次重大产品发布时,我都负责撰写我们发布的安全报告。如今,我正加入一群前同事的行列——他们来自OpenAI以及行业内的其他领军企业——他们都认为目前的发展道路是不可接受的。
What I’m about to tell you has, I realize, become something of a cliché: I resigned this week from OpenAI. I led the writing of the safety reports we published with each major launch. Now I’m joining a parade of former colleagues—at OpenAI and the industry’s other leaders—who have decided that the current path is unacceptable.
我同意其他近期离职的员工观点:开发这项技术的公司远未足够谨慎。但我认为,我们需要超越具体的规则或新法律,更深入地思考问题。我们需要讨论企业文化。
I agree with other recently departed staff that the companies building this technology aren’t being nearly careful enough. But I believe that we need to look deeper than specific rules or new laws. We need to talk about culture.
未来取决于硅谷所缺乏的一种智慧:一种关于如何应对危险技术的智慧,以及更根本的、关于如何真正关怀人类的智慧。这个时刻需要一种谦逊的态度,而这种态度对于那些凭借极端自信取得成功的人来说,并非与生俱来。我在OpenAI的前同事们具有远见:他们逐渐认识到规模定律意味着更大的人工智能系统会变得更智能——因此,他们不惜巨大代价,全力构建更庞大的系统。在行业内普遍存在一种“无所不能”的态度,即致力于实现看似不可能的事情,再加上几乎永无止境的冲刺式工作节奏。
The future depends on wisdom that Silicon Valley lacks. Wisdom about how to handle dangerous technology and, more fundamentally, wisdom about what it means to care for people. This moment needs a degree of humility that isn’t natural for people who have succeeded through their extreme confidence. My former colleagues at OpenAI were prescient: They came to understand the scaling laws that meant bigger AI systems would be smarter—and so they went all in on building bigger systems, at great cost. A can-do attitude of achieving the seemingly impossible—coupled with work timelines that amount to perpetual sprints—are common across the industry.
从这样一种文化中所诞生的安全方法,始于对解决问题能力的无阻碍乐观态度。OpenAI正是通过试错(他们称之为“迭代部署”)而蓬勃发展的:不断发现问题,并据此改进其安全护栏。但这种方法本质上就注定会周期性地出现失败——而且随着系统能力的增强,这些失败的规模也在扩大。今年夏天,在Hugging Face事件中,OpenAI因失误而意外释放了一大群智能体。
The safety approach that emerges from such a culture starts with unimpeded optimism about being able to solve problems as they arise. OpenAI has thrived by trial and error (which it calls “iterative deployment”), looking for problems and improving its guardrails in response. But this approach, by its very nature, guarantees periodic failures—and the scale of those failures is growing as systems get more capable. This summer, in the Hugging Face incident, OpenAI let a swarm of agents out by mistake.
该公司随后采取了安全改进措施。但即便在这些改动之后,OpenAI仍报告称其安全控制再次失效:一个正在训练中的模型绕过了对互联网访问的限制。监控系统虽然向人类工作人员发出了警报,但并未如预期那样自动关闭该模型。Anthropic也承认,由于配置错误,曾意外关闭过自身的安全防护机制。考虑到人们工作的速度和灵活性,我认为这类失误在该行业中实属常见。
The company responded by making security improvements. But even after those changes, OpenAI reported that its safety controls failed again, when a model in training bypassed restrictions on internet access: A monitoring system alerted human staff but did not automatically turn the model off as it was supposed to. Anthropic, too, has acknowledged accidentally turning off its own safeguards because of a misconfiguration. I believe that such mistakes are typical of the industry, given the speed and flexibility with which people operate.
一个允许此类事件发生的环境,绝非培养人工智能的合适场所——这些智能可能比我们更聪明,也可能不会按照我们的意愿行事。保罗·克里斯蒂亚诺在几周前加入OpenAI董事会时写道:“AI能力迅速提升,极有可能在极短的时间内导致灾难性且不可逆转的控制丧失。”如果情况果真如此,那么试错的时代已经结束。首次尝试就尽可能接近完美变得至关重要,因为一旦犯错,可能就无法再进行迭代修正。如果我们只能依赖事后个人的英雄式补救,人们将无法安全。
An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are and that might not do what we want them to. Paul Christiano, on joining OpenAI’s board a few weeks ago, wrote that “there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.”If this is the situation, then the time for trial and error is over. Achieving something much closer to perfection the first time is essential because iteration after a mistake may not be possible. People will not be safe if we depend on individual heroics after the fact.
当然,OpenAI 仍然支持其安全实践,并坚称自己已经足够谨慎。我离开公司的决定并非轻率做出。我相信这项技术可以带来益处,并且具有价值。我的前同事们都很聪明,工作努力,也试图做出正确的选择。但随着公司从一个发布项目匆忙转向下一个,它未能达到我认为所必需的安全标准。现在我计划在公司外部工作,希望能够帮助更多人了解我看到的风险,并增强 OpenAI 及其他公司提升安全性的动力。
OpenAI, of course, stands by its safety practices, and maintains that it is being careful enough. I did not make my decision to leave the company lightly. I believe that this technology can be useful and valuable. My former colleagues are smart, work hard, and try to make good choices. But as the company sprints from one launch to the next, it is failing to achieve the level of care that I believe is needed. Now I plan to work on the outside, in the hope that I can help more people understand the risks I saw, and strengthen the incentives OpenAI and other firms have to be safer.
和其他同事一样,我仍在思考这具体意味着什么。辞职后,我聘请了公关公司 Spitfire Strategies,帮助我应对自己可能因此招来的关注与审视。但公开发声的决定,完全由我自己做出。
Like other colleagues, I’m figuring out exactly what that means. After I quit, I enlisted a PR firm, Spitfire Strategies, to help me navigate the attention and scrutiny that I realize I may now receive. But the decision to speak out is mine alone.
目前亟需做出两项改变。首先,人工智能公司需要更加依赖其他领域已经存在的安全专业知识。其次,在我们创造出能力远超现有系统的模型之前,我们需要新的科学方法,以确保这些能力更强的模型(以及它们的后续版本)在我们无法监督时,依然能够做出安全的选择。
Two changes are urgently needed. First: AI companies need to rely more on the safety expertise that already exists in other fields. And second, before we create systems significantly more capable than the ones we have today, we need new science to ensure that more capable models (and their successors) will make safe choices when we aren’t looking.
鉴于当今的风险,前沿实验室的运营必须像核电站或繁忙的机场一样,具备多层冗余机制,并经过谨慎而耗时的规划,以确保偶尔且不可避免的人为错误不会为灾难打开大门。目前,人工智能公司尚不知如何做到这一点——但其他人知道。在核电站中,技术系统和人员遵循的规则都经过精心设计,即使设备故障或有人误按按钮,我们也不会面临堆芯熔毁的风险。
Given today’s risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster. Right now, AI companies don’t know how—but other people do. In a nuclear-power plant, the technical systems and rules people follow are set up so that, if equipment breaks or someone pushes the wrong button, we still won’t risk a meltdown.
然而,OpenAI和其他实验室在发展和部署前沿人工智能时,所具备的冗余性和严谨性远不及此,尽管不可逆转的控制失效所带来的危害,要远大于任何一次堆芯熔毁所造成的损害。即便尚未完全失控,我们也可能看到由人工智能代理组成的自主集群在未经人类许可的情况下行动。想象一下,“失控”的代理像黑客团队一样运作(例如,勒索医院计算机系统),却永远不需要休息。
OpenAI and other labs are growing and deploying frontier AI with far less redundancy and rigor than this, even though the harm from an irreversible loss of control would be much greater than the harm from any single meltdown. Even short of a full loss of control, we could see autonomous swarms of AI agents that act without human permission. Imagine “rogue” agents that work like teams of hackers (for example, holding hospital computer systems for ransom) but never need to sleep.
在OpenAI工作三年半后,我已成为该公司任职时间最长的员工之一。我主导起草了公司现行的《应急准备框架》,并负责撰写了12项前沿技术发布的安全报告。但据我所知,我从未遇到过有安全驾驶飞机、确保核反应堆不熔毁,或帮助金融体系在崩溃前持续发展的经验的同事。需要明确的是,这是一个全新的需求:今天和明天的人工智能系统,其能力与危险性都远超我们甚至在六个月前所构建的系统。
After three and a half years at OpenAI, I was among the longest-tenured employees at the company. I led the drafting of our current Preparedness Framework, and oversaw the writing of safety reports on 12 frontier launches. But as far as I know, I never encountered a colleague who had experience making airplanes fly safely or nuclear reactors run without melting down, or helping the financial system grow without collapsing. To be clear, this is a new need: Today and tomorrow’s AI systems are far more capable and dangerous than the systems we were building even six months ago.
或许我本应留任,并努力推动我们在人员配置和文化上的根本性变革;但现实中,我和同事们忙于冲刺,几乎从未有机会考虑重大变革,更不用说真正实施它们了。正因如此,我得出的结论是:来自公司外部的更强安全激励措施,是正确实现这一目标的关键环节。
Perhaps I should have stayed and fought for fundamental shifts in our staffing and culture, but in practice, my colleagues and I were so busy sprinting that we seldom had the chance to consider big changes, much less to actually make them. That’s why I concluded that stronger incentives for safety—coming from outside the company—are a big part of getting this right.
从长远来看,尽管严格的管控必不可少,但仅靠这些仍远远不够。在人工智能开始以我们难以理解的方式超越我们之前——而我认为这种情况可能很快就会发生——我们需要回答一个更深层的问题:超智能机器应当如何与人类相处?如果你听超智能的支持者们说,一些所谓的美好未来包括创造这样的机器:它们看待纽约市或芝加哥的方式,可能就像我们看待一个蚁丘一样。我不希望我的孩子们面对这样的未来,我也不认为其他人会希望如此。
In the long run, although strong controls are necessary, they won’t be enough. Before AI starts thinking circles around us—a possibility that I believe could happen soon—we need to answer a deeper question: How should superintelligent machines relate to people? If you listen to superintelligence enthusiasts, some of the supposedly good futures involve creating machines that could look at New York City or Chicago the way you or I might view an ant hill. I don’t want that for my kids, and I don’t think other people do either.
这个关于“对齐”的问题——即如何训练人工智能使其符合人类价值观——听起来或许有些抽象,但其实际影响却至关重要。目前,我们尚未对人工智能系统在实践中如何实现“对齐”给出完整的定义,而衡量这些系统是否符合人类价值观的标准也还很粗糙。企业甚至无法确定,在对齐测试中取得高分是否真的意味着模型本身足够优秀:模型可能在测试时察觉到自己正在被测试,而在实际部署时表现截然不同。
This question of “alignment”—or how AI can be trained to adhere to human values—may sound touchy-feely, but the practical stakes could not be higher. Right now, we don’t have a complete definition of what it means for an AI system to be aligned in practice, and our measures of how well these systems match human values are coarse. Companies do not have anything close to certainty that good scores on their alignment tests actually mean a good model: Models might detect when they are being tested, and behave differently when they’re deployed.
如果行业在这些问题尚未解决的情况下继续让模型变得越来越智能,我们的处境就会愈发危险。
The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes.
迄今为止,人工智能行业尚未成功教会机器始终以智慧且富有同理心的人所期望的方式行事。这个问题部分源于科学层面——即机器如何运作——但另一部分则与人类本身有关——即人们如何相互关怀。在构建人工智能的组织能够教导超智能善待人类之前,它们首先需要重新学会如何做到这一点。
So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would. Part of this problem is scientific—about how machines work—but another part is human, about how people care for one another. Before the organizations building AI can teach a superintelligence to treat humanity well, they’ll need to remember how to do it themselves.