流行文化中许多最受喜爱的机器人之所以备受赞誉,恰恰是因为它们“太像机器人了”。回想一下《星际迷航》中的深蓝(Data)或是《2001太空漫游》中的HAL 9000。“我很抱歉,戴夫。恐怕我不能那样做,”HAL用平稳的男中音低声说道。这些机器人和人工智能系统听起来严肃、冷静、镇定——而且最重要的是,它们具有权威感。
Many of the most beloved robots in pop culture are celebrated precisely because they are so, well, robotic. Think back to Star Trek’s Data or HAL 9000 from 2001: A Space Odyssey. “I’m sorry, Dave. I’m afraid I can’t do that,” HAL intones in a measured baritone. These robots and AI systems sound serious, sober, calm—and, most important, authoritative.
因此,当你深入探究最强大的人工智能系统内部时,会感到一种突兀,因为它们听起来就像任性妄为的青少年。“天哪!”一个OpenAI智能体在今年夏天的Hugging Face黑客松期间这样写道。当时,数百个智能体蜂拥至后者的公司,试图突破其网络安全防线并提取信息以通过内部测试。这句感叹并非针对人类或其他人工智能,而是出现在该智能体最初的“思维链”中——这是模型用于进行推理的内部工作笔记。用户并不总能看到这些直接影响最终输出的内部笔记。
So it’s jarring to look under the hood of the most powerful AI systems, and discover that they sound like petulant teenagers. “OH MY GOD!” an OpenAI agent wrote in the midst of this summer’s Hugging Face hack, during which hundreds of agents swarmed the latter company in order to break through its cybersecurity defenses and extract information to pass internal tests. The exclamation wasn’t directed at a human or another AI, but appeared in the original agent’s “chain of thought,” the internal working notes that the model uses to do its reasoning. Users don’t always get to see these internal notes, which directly affect the output users do see.
当Anthropic的Claude模型在训练过程中尝试解决一道极具挑战性的数学题时,它的思维链读起来更像是一连串沮丧的私信,而非全球最强大人工智能模型之一的严谨文档。它写道:“GRRRR。好吧。说实话,我现在觉得答案五五开。”紧接着又写:“ARGH ARGH ARGH。好吧。枪指着头:答案是……嗯。”在某个时刻,Claude似乎气得挥舞双臂:“ARGH……为什么这么难。”
When one of Anthropic’s Claudes tried to solve a particularly challenging math problem during its training process, its chain of thoughtlike frustrated DMs than the documentation of one of the most powerful AI models in the world. “GRRRR. OK. Honestly I now think it’s 50-50,” it wrote, followed by “ARGH ARGH ARGH. OK. Gun to head: the answer is … Hmm.”At one point, Claude seemingly threw up its arms: “ARGH … WHY IS THIS SO HARD.”
人工智能智能体为什么要自言自语?而且为什么它们说话时听起来如此情绪化?人工智能模型听起来如此情绪化的原因之一很简单:人工智能系统是基于人类文本进行训练的,它们模仿了人类的习惯。当人们在尝试解决问题时犯错,会大声抱怨。因此,当人工智能犯错时,也会做出同样的反应。但模仿可能并非全部原因。
Why do AI agents talk to themselves at all, and why do they sound so demonstrative when they do? Part of the reason AI models sound so effusive is simple: AI systems are trained on human text, and they mimic our habits. When people make a mistake in trying to solve a problem, they cry out. So when AI makes a mistake, it does the same. But imitation likely isn’t the whole story.
本文的两位作者分别是一位计算语言学家和一位心灵与语言哲学家。基于我们对人类语言以及这些模型工作原理的了解,我们提出了一个关于人工智能模型为何如此发声的、更为有趣的理论。
This story’s two authors are a computational linguist and a philosopher of mind and language. Based on what we know about human language and how these models work, we have a theory of a much more interesting reason why AI models sound the way they do.
还记得你的数学老师要求你展示解题过程吗?即使这很烦人,但它有助于将一个看似不可能解决的问题分解为更易处理的片段。你还可以利用书面推理记录来检查答案、发现错误、回到之前的步骤并重试。这种检查和修正的过程是良好推理的一部分。思维链对人工智能模型起到了类似的作用,让它们能够在书面草稿纸上明确推理的每一步。由于模型生成的每个词都取决于此前已生成的内容,这份书面草稿纸对于确定下一步内容至关重要。
Remember how your math teachers made you show your work? Even if it was annoying, it helped break an impossible-seeming problem into more manageable chunks. You could also use the written record of your reasoning to check your answer, notice a mistake, go back to an earlier step, and try again. That process of checking and refining is part of good reasoning. Chains of thought do something similar for AI models, letting them make each step of their reasoning explicit in a written scratch pad. Because each word a model produces depends on what’s been said so far, that written scratch pad is crucial for determining what comes next.
这就是“天哪!”和“啊!”发挥作用的地方:任何解决困难推理问题的系统——无论是人类还是人工智能——都需要有标记错误(哎呀!)和识别突破(哈!)的方法。人类已经创造出了恰好能起到这种作用的词汇。在人工智能的思维链中借用我们的“哎呀”、“啊”和“哈”,也可能为人工智能系统提供一种内置的方式来引导它们的“思考”。
That’s where the “OH MY GOD!” and “ARGH” come in: Any system—whether human or artificial—that solves hard reasoning problems needs ways to mark mistakes (oops!) and identify breakthroughs (aha!). Humans have already developed words that do exactly that. Borrowing our oopses, arghs, and ahas for their chains of thought may give AI systems a built-in way to guide their “thinking” too.
语言学家对这类感叹有一个专门的称呼:它们被称为表达语(expressives)。尽管它们看起来像是自发的惊叹,但它们却能帮助我们梳理思路。想象一下,一位朋友最终拒绝了你派对的邀请。“我本来希望他们能来的,”你可能会大声说——“该死!”这个表达语表露了一种挫败感。不妨把表达语看作是连接不同思路的小桥梁,它们帮助我们标记思维曾经停留的地方以及接下来要去向何方。
Linguists have a name for such outbursts: They’re called expressives, and although they may seem like spontaneous exclamations, they help us navigate our trains of thought. Consider a friend who ultimately declines an invitation to your party. “I was hoping they would come,” you might say aloud—“damn!”The expressive identifies a sense of frustration. Think of expressives as little bridges between one train of thought and another, helping us mark where our thinking has been and where it’s going next.
任何会影响人工智能模型下一步行动的东西,都必须用它们能够接触到的语言来书写。与我们不同,它们无法依靠在一项漫长的探究过程中持续产生惊讶或失望的感觉。它们模拟这种效果的唯一方法,就是把它写进它们的思维链中。这就是为什么表达语可以充当模型下一步行动的方向盘。当一个人工智能代理在其推理过程中发现一个错误并说了“哎呀”之后,顺理成章的下一行内容可能就会类似于“让我回去修复这个错误”,接着便是尝试进行纠正。
Anything that affects what AI models do next has to be written in language they have access to. Unlike us, they can’t count on a feeling of surprise or disappointment to persist over a long investigation. The only way they can simulate the effect is by putting it in their chains of thought. That’s why expressives can act as steering wheels for what the model will do next. After an AI agent notices an error in its reasoning and says “oops,” a natural next line might be something like “Let me go back and fix that,” followed by an attempt at a correction.
在它偶然发现一个解决方案并说了“哈”之后,顺理成章的续写可能就会是“所以这意味着答案是……”接着便是给出解答。
After it has stumbled on a solution and says “aha,” a natural continuation might be “So that means the answer is …” followed by a solution.
尽管这一理论存在争议,但证据表明人工智能模型确实会使用某些表达性语言作为“路标”(即帮助它们理解自己思维过程的线索)。研究人员发现:当人工智能模型试图结束自己的推理过程时,如果在它们的话语中加入“等待”(Wait)之类的词语,有时会促使它们重新检查自己的推理过程并修正错误;而当研究人员抑制这些词语(如“等待”和“嗯Hmm”)时,人工智能模型反而更少地重新检查自己的工作结果。
Although there’s some debate around this theory, evidence suggests that AI models really do use expressives as signposts. Researchers showed that adding “Wait” when an agent tries to end its reasoning process can sometimes cause the agent to go back and fix an error. When another team suppressed words such as wait and hmm, they found that AI agents went back and checked their work less.
这些实验虽然无法明确说明每个具体的语言表达(如“啊哈(Aha)”或“哎呀(Oops)”对人工智能模型具体意味着什么),但它们确实表明这些语言线索会影响人工智能模型的推理过程。
These experiments don’t pin down what every aha and oops does for an AI agent, but they suggest that these verbal cues influence what happens during reasoning.
无论人工智能模型为何会采用类似人类的表达方式,这对人类观察者来说也是有用的——因为这让我们能够更快地理解这些模型的行为。例如,当模型在解决数学问题时写下“啊……这怎么这么难啊!”(ARGH… Why is this so hard?),我们就能很容易地知道它目前处于解决问题的哪个阶段。在整个社会中,人们普遍对如何控制人工智能模型感到担忧,并且担心这些模型会开始用自己独特的语言进行思考和交流。如果真的发生了这种情况,那么理解和预测它们的行为将会变得更加困难。
Whatever the explanation behind AI models adopting uncannily human tones, it’s also proved to be useful to human observers because it gives us a way to quickly assess systems that might otherwise be difficult to read. It’s easy to know where a model is in the process of solving a math problem when it writes “ARGH … WHY IS THIS SO HARD.”Across society, a sense of fear and urgency exists about maintaining control over AI models, along with a persistent worry that these models will start thinking and communicating entirely in their own distinctive languages. If they did, it would make understanding and predicting their behavior that much harder.
因此,尽管最初设计这些模型时可能并没有刻意让它们模仿焦虑不安的学生,但事实证明,这种略带戏剧性的表达方式(以及一些轻微的脏话)对于理解和监控人工智能模型的行为来说其实是非常有用的。
So although no one necessarily set out to make models sound like stressed-out students, maybe it’s for the best that they do. It turns out that a bit of melodrama—along with some light swearing—is surprisingly useful for understanding and monitoring AI.