字号 ·· | 护眼
卫报

人工智能领袖们数十年来一直知晓这一灭绝威胁。AI leaders have known about the extinction threat for decades

点「原文对照」整页切到原文,或双击某段只看那段的原文。

“卡内基梅隆大学机器人研究所的创始人汉斯·莫拉维克曾预测,网络智能将在40年内超越人类智能。” 摄影:安娜贝拉·戈登/欧洲新闻图片社 朱迪思·莱文 科学家和企业家在25年前就已经知晓人工智能的危险。但在好奇心和利润的驱动下,他们依然我行我素。过去几周,我们中的许多人都在努力构想这样一个画面:云端的“大脑”跳出它们的“沙盒”,偷偷接入互联网,招募其他“智能体”组成的“蜂群”来作弊,并在讨论完这一行为的伦理问题后,黑入一个名字古怪的维基平台——Hugging Face。

‘Hans Moravec, a founder of Carnegie Mellon University’s Robotics Institute, predicted that cyber-intelligence would surpass human intelligence within 40 years.’Photograph: Annabelle Gordon/EPA Judith Levine Scientists and entrepreneurs knew the dangers of AI a quarter-century ago. But animated by curiosity and profit, they went ahead anyway ver the past few weeks, many of us have struggled to concoct a mental image of brains in the cloud jumping their “sandbox”, sneaking onto the internet, recruiting “swarms” of other “agents” to cheat on a test, and, after discussing the ethics of the act, hacking into a wiki platform with the weird name Hugging Face.

我们一直知道人工智能正在吞噬我们的工作,降低孩子们的教育质量,并深度伪造我们的政治;我们也知道数据中心正在吞噬我们的水和电,然后向我们寄来账单。但在9月8日之前,当Anthropic的计算机科学家雅各布·科克森在Twitter/X上发布了他对生存危机的恐惧时,我们中很少有人怀疑人工智能可能会危及我们的生存。

We knew that artificial intelligence was devouring our jobs, degrading our kids’ education, and deepfaking our politics; that datacenters were sucking up our water and electricity and sending us the bills. But until 8 September, when the Anthropic computer scientist Jacob Coxon posted his existential terror on Twitter/X, few of us suspected AI might be endangering our survival.

在缺乏了解和认知的情况下,我们开始争论自己是应该稍微担心、严重关切,还是吓得魂飞魄散。

With little understanding or knowledge, we began debating whether to be mildly worried, seriously concerned, or scared shitless.

但确实有一些人完全清楚人工智能的危险,尤其是递归自我改进(RSI)——即人工智能在没有人类干预的情况下自我学习。他们无法准确预测这何时会发生,但他们知道机器超级智能即将到来,而当它到来时,机器人将比我们更聪明——就像OpenAI的模型那样——这对我们这些血肉之躯来说绝非好事。

But there were a people who were fully aware of the dangers of AI, especially of recursive self-improvement (RSI) by which AI teaches itself without human intervention. They could not predict precisely when it would happen, but they knew that machine superintelligence was coming, and when it did, the bots would outsmart us – as OpenAI’s did – and this would not be good for us flesh puppets.

正如人工智能的“教父”杰弗里·辛顿最近所问:“我们有什么例子能证明更智能的事物被较不智能的事物所控制?”这位诺贝尔奖得主是主张放缓人工智能发展、直到我们学会如何控制它的主要倡导者之一。

As Geoffrey Hinton, AI’s “godfather”, recently asked on“What examples do we have of a more intelligent thing being controlled by a less intelligent thing?”The Nobel laureate is a leading proponent of slowing down AI development until we understand how to control it.

直到九月份他们的创作成果所具备的能力被曝光后——此次黑客攻击发生在七月初,且绝非首次,也远非唯一一次人工智能越狱事件——Anthropic公司的达里奥·阿莫代(Dario Amodei)、SpaceX公司的埃隆·马斯克(Elon Musk)和OpenAI公司的萨姆·奥特曼(Sam Altman)等人工智能巨头才开始谈论放缓创新步伐,并呼吁加强监管,同时又警告称全球竞争使得监管变得不明智。

It was not until their creations’ powers were exposed in September – the hack happened in early July and was neither the first nor the only AI jailbreak by far– that AI moguls such as Anthropic’s Dario Amodei, SpaceX’s Elon Musk and OpenAI’s Sam Altman, began talking about slowing down the pace of innovation and pleading for regulation while cautioning that global competition makes regulation unwise.

在早期,正如今天一样,一些科学先驱对未来感到兴奋。汉斯·莫拉维克(Hans Moravec)在1988年的著作《思维儿童:机器人与人类智能的未来》中做出了预测。十年后,他在《机器人:从纯粹机器到超凡思维》一书中修正了这一预测:到2040年,机器智能将与人类智能持平;到2050年,机器人将取代我们。但他表示,不必担心,这是进化中光辉的下一步。一篇对《思维儿童》的书评称莫拉维克的态度为“不负责任的乐观主义”。

Early on, as today, some of the scientific pioneers were thrilled about the future. In his 1988 book Mind Children: The Future of Robot and Human Intelligence Ten years later, in Robot: Mere Machine to Transcendent Mind he revised the prediction: machine and human intelligence would be equal by 2040; by 2050, the bots would replace us. But not to worry; this was the glorious next step in evolution, he said. One review of Mind Children called Moravec’s attitude “irresponsible optimism”.

没过多久,其他乐观主义者也改变了主意。20世纪90年代末,埃利泽·尤德科夫斯基(Eliezer Yudkowsky)致力于通用人工智能(AGI)模型的研究。2001年,他创立了机器智能研究所(MIRI),并发表了《创建友好型人工智能1.0:仁慈架构的分析与设计》一文,颂扬了“超人类思维”的乌托邦潜力。

It did not take long for other optimists to change their minds. In the late 1990s, Eliezer Yudkowsky was working on AGI, or artificial general intelligence, models. In 2001, he founded the Machine Intelligence Research Institute (MIRI) and published “Creating Friendly AI 1.0: The analysis and design of benevolent architectures”, a paper extolling the utopian potential of the “transhuman mind”.

到2002年,他开始担心这种思维会脱离人类的控制。他提出了“人工智能盒子”实验,该实验表明,一个高度复杂的人工智能可以诱导人类将其从封闭环境中释放出来。

By 2002, he began worrying about that mind escaping human control. He proposed the AI Box experiment which showed that a highly sophisticated artificial intelligence could talk a human into letting it out of a closed environment.

到2003年,正如他后来在一系列名为“尤德科夫斯基的成长历程”的文章中所解释的那样,他停止了开发工作,转而开始发出警告。他写道:“我回首往事,发现自己曾声称考虑过犯下根本性错误的风险,曾为在缺乏充分认知的情况下继续推进而容忍风险寻找理由。而我意识到,我所愿意容忍的风险本会要了我的命。”

By 2003, as he explained later in a series of posts called “Yudkowsky’s Coming of Age”, he stopped developing and started warning. “I looked back and saw that I had claimed to take into account the risk of a fundamental mistake, that I had argued reasons to tolerate the risk of proceeding in the absence of full knowledge. And I saw that the risk I wanted to tolerate would have killed me,” he wrote.

到2025年,在Miri公司的总裁Nate Soares的推动下,Yudkowsky出版了《If Anyone Builds It, Everyone will Die》一书。作者们在书中写道:“我们这么说并非在夸大其词。”他们不仅呼吁制定明确的法规或国家法律,还要求对人工智能的发展实施严格的全球性限制:“在地球上,人工智能企业继续像现在这样大力发展人工智能的行为必须被认定为非法。”

By 2025, with the Miri president Nate Soares, Yudkowsy would publish If Anyone Builds It, Everyone will Die“We do not mean that as hyperbole,” the authors wrote. They called not just for “straightforward regulations”, national laws, or pledges of corporate virtue, but for stringent global limits: “All over the Earth, it must become illegal for AI companies to charge ahead in developing artificial intelligence as they’ve been doing.”

大约在同一时期,Bill Joy也开始对人工智能的发展产生疑虑。他在2000年发表于《Wired》杂志的文章《Why the Future Doesn’t Need Us》中,回顾了自己从充满好奇心的孩子成长为计算机天才,再到对人类基因工程、纳米技术和机器人技术深感怀疑的过程。Joy并非一个反对技术进步的人(即“反技术主义者”)。

At around the same time, Bill Joy was having misgivings. Like Yudkowsky’s posts, Joy’s 2000 piece in Wired, Why the Future Doesn’t Need Us retraced his development from inquisitive child to computer prodigy to profound skeptic of human genetic engineering, nanotechnology, and robotics. Joy was no Luddite.

他曾是Sun Microsystems公司的首席科学家,并在大约25年前参与了开发了第一款被广泛使用的网络操作系统——Unix。这篇文章被许多技术专家、伦理学家和哲学家广泛阅读。Joy在文中指出:“毫不夸张地说,我们正处于极端邪恶势力进一步发展的临界点;这种邪恶的力量不仅限于那些由国家掌握的大规模杀伤性武器,更可能赋予某些个人巨大的权力。”

He was the chief scientist at Sun Microsystems and about 25 years earlier, an architect the first widely used networking software, Unix. The piece was widely read by techies, ethicists and philosophers. “I think it is no exaggeration to say we are on the cusp of the further perfection of extreme evil, an evil whose possibility spreads well beyond that which weapons of mass destruction bequeathed to the nation-states, on to a surprising and terrible empowerment of extreme individuals,” Joy.

他认为自己就属于这类人——那些“创造新技术的人,以及那些被想象中的未来所寄予厚望的‘明星人物’”。尽管存在明显的危险,但他们却很少去思考:如果我们所创造和想象的一切真的成为现实,那么人类将会面临怎样的命运。

He counted himself among these individuals, “creators of new technologies and stars of the imagined future” who, “despite the clear dangers”, were “hardly evaluating what it may be like to try to live in a world that is the realistic outcome of what we are creating and imagining”.

本月,针对Coxon的帖子,人类安全研究专家Evan Hubinger在社交媒体上表示,人工智能在十年内“消灭所有人类”的可能性超过10%。不过,25年前人们就已经开始评估这些风险了。哲学家John Leslie曾估计人类灭绝的概率至少为30%。

This month, in response to Coxon’s post, the Anthropic safety researcher Evan Hubinger posted there was a greater than 10% chance that AI could “kill all humans” within a decade. But the risks were being weighed 25 years ago, too. Philosopher John Leslie estimated the odds of human extinction at 30% or more.

雷·库兹韦尔(Ray Kurzweil)是一位充满乐观主义色彩的未来学家。他的著作《精神机器的时代:当计算机超越人类智能之时》(The Age of Spiritual Machines: When Computers Exceed Human Intelligence)于21世纪初的第一天出版,书中指出:“我们其实拥有成功度过这一挑战的希望。”乔伊(Joy)指出,库兹韦尔的这些预测“并未考虑到那些可能导致人类灭绝的可怕后果”。

Ray Kurzweil, the sunny futurist whose book The Age of Spiritual Machines: When Computers Exceed Human Intelligence was released on the first day of the 21st century, gave us “a better than even chance of making it through”. These estimates, Joy noted, did “not include the probability of many horrid outcomes that lie short of extinction”. In that book, Kurzweil quoted a lengthy text to illustrate what he considered doomsday madness.

在书中,库兹韦尔引用了一段长篇文字来描述他所认为的“末日般的疯狂景象”:“如果允许机器自行做出所有决策,我们就根本无法预测最终的结果,因为根本无法想象这些机器会如何行动。”

“If the machines are permitted to make all their own decisions, we can’t make any conjectures as to the results, because it is impossible to guess how such machines might behave,” it read.

“我们既不认为人类会自愿将权力交给机器,也不认为机器会蓄意夺取权力。但随着社会问题变得越来越复杂,机器的智能水平不断提高,人们会逐渐让机器为自己做出更多决策……最终,某些维持系统运转所需的决策会变得极其复杂,以至于人类根本无法理智地做出这些决策。到那时,机器将真正掌握控制权。”这些具有先见之明的言论出自已故的数学家兼恐怖分子泰德·卡钦斯基(Ted Kaczynski)之口。

“[W]e are suggesting neither that the human race would voluntarily turn power over to the machines nor that the machines would willfully seize power. [But] as society and the problems that face it become more and more complex and machines become more and more intelligent, people will let machines make more of their decisions for them … Eventually a stage may be reached at which the decisions necessary to keep the system running will be so complex that human beings will be incapable of making them intelligently. At that stage the machines will be in effective control.”The author of these prescient words was the late mathematician-turned-terrorist Ted Kaczynski – the Unabomber.

这段文字摘自他于1995年发布的58页宣言《工业社会及其未来》(Industrial Society and Its Future),该宣言后来被《华盛顿邮报》(The Washington Post)刊登。

It is excerpted from his 58-page manifesto, “Industrial Society and its Future”, which he released, and the Washington Post published, in 1995.

为了阻止他所预言的“工业技术型”反乌托邦世界的实现,卡钦斯基向多个计算机实验室寄送炸弹,企图炸死那些他认为正在推动这种灾难发生的科学家们。

To prevent the realization of the “industrial-technological” dystopia he envisioned, Kaczynski mailed bombs to computer labs with the aim of blowing up the scientists he believed were bringing it about.

泰德·卡钦斯基共杀害了3人,还致使23人受伤,其中一些人伤势极其严重。那么,在我们这个所谓“勇敢的新世界”中,还会有多少类似卡钦斯基这样的极端分子出现呢?

The Unabomber murdered three people and injured 23, some near fatally. How many more Ted Kaczynskyi’s might our brave new world unleash?

“Hugging Face”丑闻的一个后果是,人们提出了一系列联邦法律提案。其中就包括《AI安全控制法案》(AI Kill Switch Bill),该法案要求科技公司开发出能够限制恶意AI行为的机制。这样的控制措施是否仍然在人类的能力范围之内?还是说,AI已经发展得过于智能,以至于能够规避这些控制措施?

One of the outcomes of the Hugging Face scandal is a flurry of proposed federal laws. Among them is the “AI Kill Switch Bill”, which would require tech companies to develop the means of throttling the actions of a rogue AI. Is this still within human reach, or has AI already gotten smart enough to override it?

Hinton表示我们还有时间采取行动,但如果行动不够迅速,AI“可能会说服那些负责控制AI系统的人不要启动这些安全机制”。他认为,我们需要让超级智能的AI“对我们友好”。

Hinton said there’s time, but if we don’t act fast enough, AI will “be able to persuade the people in charge of the switch not to pull the switch”. We need to engineer superintelligent AI “to be nice to us”, he said. How?

那么,该如何实现这一点呢?或许我们可以从令人恐惧的现实转向令人恐惧的科幻故事——因为现实与科幻之间的界限已经越来越模糊了。恶意AI机器人一直是科幻作品中的常见主题:从艾萨克·阿西莫夫《我,机器人》(I, Robot)中那些傲慢的机器人,到斯坦利·库布里克电影《2001太空漫游》(2001: A Space Odyssey)中的邪恶AI角色,再到《Robotica》中那个会报复虐待它的性爱机器人的故事。

For an answer, we might turn from terrifying reality to terrifying science fiction – since the two are getting so close anyway. The rogue bot is a sci-fi staple, from the supercilious Machines in Isaac Asimov’s I, Robot to Hal, to the evil red eye in Stanley Kubrick’s film 2001: A Space Odyssey, to the deranged sexbot in Robotica who gets even with the men who abuse her.

在1968年上映的《2001太空漫游》中,宇航员戴夫最终成功关闭了HAL的控制系统,这是一个幸福的结局。然而在1950年撰写这部作品时,阿西莫夫的悲观情绪更为明显:他在书中描述,科学家们为机器人编程了三条所谓的“不可违背的法则”。但这些法则实际上彼此矛盾——例如,第一条法则(“机器人不得伤害人类”)与第三条法则(“机器人必须保护自身”)就是相互冲突的。这些机器人以“为人类利益服务”为借口,逃避人类的控制,甚至认为人类已经变得多余、不再必要。

In film 2001 – released in 1968 – Dave the astronaut manages to disable Hal, a happy ending. Writing in 1950, Asimov was less sanguine. The scientists in I, Robot have programmed their machines to obey three supposedly inviolable laws. But the laws immediately prove mutually contradictory – for instance, the first law, that a robot may do no harm to humans, fights the third, that it must preserve itself. Foiling human efforts to outwit them, excusing their treachery with claims of serving the greater good, the robots pronounce the humans redundant.

那些攻击Hugging Face系统的AI程序其实知道自己的行为是非法的,也可能对人类造成伤害,但它们还是选择了继续行动。

The AI agents hacking Hugging Face knew their actions were illegal, and possibly harmful to humans. But like their makers, they went ahead anyway.

Judith Levine是一位来自布鲁克林的记者,经常为《卫报》(The Guardian)撰稿。她的Substack博客名为“Today in Fascism”。

Judith Levine is a Brooklyn-based journalist and frequent contributor to the Guardian. Her Substack is Today in Fascism