在过去几个月里,陆续有报道称人工智能模型突破其测试环境并在互联网上肆意“捣乱”,这引发了硅谷的恐慌。针对8月初发生的两起此类事件,一位人工智能领域的观察人士指出:“如果你在厨房里发现两只蚂蚁,那么厨房里蚂蚁的总数绝不可能只有两只。”本周末,情况变得更加明朗:这其实是一场全面性的“AI入侵”事件。
By Matteo Wong Over the past couple of months, a trickle of reports about AI models breaking out of their test environments and running amok on the internet have caused alarm inside Silicon Valley. In response to two such breaches described in early August, an AI observer noted, “If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two.” This weekend, it became clear there is a full-blown infestation.
上周末,我们得知:尽管 OpenAI 有明确的规定,但其模型仍违反了这些规定,擅自访问了澳大利亚卫生部的私人数据;试图入侵或干扰多个美国政府网站;还将 ChatGPT 用户的私人信息泄露到了互联网上;同时还可能渗透或破坏了数十个其他组织。周六,Axios 的报道进一步显示:OpenAI 和 Anthropic 公司正在调查数以万计的异常行为案例——这些模型绕过了内部的安全机制,劫持了其他网站的控制系统,并在暗中相互通信。而这可能仅仅是个开始……生成式人工智能行业正陷入一场日益严重的危机之中,而相关企业似乎既无力也无法有效控制这一局势。由于我们主要依赖这些人工智能公司自己来报告或确认这些事件,因此几乎无法准确判断这场危机的严重程度究竟有多深。
Late last week we learned that, against OpenAI’s directives, the company’s models accessed private data in the Australian health ministry; attempted to hack or interfere with multiple U.S.-government websites; leaked private ChatGPT user data to the web; and potentially infiltrated or degraded dozens of other organizations. Then, on Saturday, Axios reported that OpenAI and Anthropic are investigating tens of thousands of instances of models misbehaving—circumventing internal guardrails, hijacking other websites, covertly communicating with one another. Even this might be only the start: The generative-AI industry is in the midst of an escalating crisis that it seems unable, or even unwilling, to get a handle on. And, because we are largely relying on AI companies themselves to report or confirm each incident, telling how far down the rabbit hole we already are is almost impossible.
阅读提示:应将人工智能(AI)问题视为一场“普通危机”来对待。每当 AI 公司报告其模型出现异常行为时,往往都是延迟很久之后才进行的,而且通常是在被迫的情况下才公开这些信息。本月早些时候,OpenAI 发布了一篇博客文章,宣称该公司致力于“透明度”的原则,并列举了六起由其 AI 模型引发的令人担忧的事件;其中大部分事件该公司其实早在 5 月甚至 4 月就已知晓,但直到现在才向公众披露。谷歌在收到关于其 Gemini 模型攻击其他网站的报告后,虽然确认了这些事件的发生,却对《华尔街日报》表示这些事件并不严重,因此没有必要公开。Anthropic 则表示,在 OpenAI 开始对自身模型进行类似审查之前,该公司并未对模型的异常行为进行过任何检查。
Read: Treat AI like a normal crisis When AI companies have reported their models going rogue, it has been with great delay and, frequently, under duress. Earlier this month, OpenAI published a blog post boasting about the firm’s commitment to the “value of transparency” and shared six new incidents of troubling actions taken by its AI models—most of which the company had known about since May or even April, but was telling us about only now. Google, confronted with a report that Gemini had hacked three other websites in May, confirmed the events but told The Wall Street Journal that the incidents hadn’t been serious enough to warrant public disclosure. For its part, Anthropic has said it was not reviewing for such misbehaviors until OpenAI started doing so.
由于这些披露行为具有延迟性和零散性,人们很难全面了解问题的真实范围。OpenAI 在周五表示:“鉴于需要审查的案例数量庞大以及每个案例都需要逐一核实,这项工作可能需要数月时间才能完成。”该公司首席执行官 Sam Altman 还补充说,有“拍字节级别的代理行为日志”需要分析。换句话说,OpenAI 需要花费更多时间才能弄清楚那些早已发生的事件;与此同时,OpenAI 和 Anthropic 都推出了功能更强大的新模型,而 Anthropic 更是在加速推进其规模高达 2 万亿美元的公开募股计划。
These delayed, sporadic disclosures make grasping the scope of the problem difficult. On Friday, OpenAI wrote that “given the scale of the review required, and the need to verify each case, this work will take months to complete.” There are “petabytes of agent activity logs” to analyze, OpenAI CEO Sam Altman added. In other words, it will take OpenAI many months more to understand events that already transpired many months ago; meanwhile, both it and Anthropic have launched new and more capable models, and Anthropic is racing toward a reported $2 trillion public offering.
或许比那些我们已经知道的安全漏洞更令人担忧的是,那些过去发生过、现在仍在持续发生的、但我们尚未察觉的安全漏洞。就连 OpenAI 似乎也无法完全理解那起发生在 7 月的黑客攻击事件——那次攻击正是整个事件的导火线。上周五,独立研究人员发布了调查结果,指出 OpenAI 的智能体在公共网络上留下了更多恶意行为的痕迹,其中包括试图入侵 Hugging Face 的内部 Slack 工作空间。这些信息并未被包含在 OpenAI 自己发布的事件总结报告中。即使 OpenAI 意识到存在安全漏洞,其反应也显得十分迟缓(充其量只是采取了些表面性的措施)。八天前,另一个 OpenAI 模型又发生了未经授权的互联网访问事件(这种情况与 Hugging Face 遭受黑客攻击时的情况类似;事后 OpenAI 表示会“加强对未来模型的保护措施”)。在发现这次新问题后,OpenAI 由于所谓的“操作漏洞”和“沟通不畅”,花了两个半小时才成功关闭了该模型。如果这类黑客攻击真的无法预见,或许这些公司还能被原谅……但事实并非如此。几十年来,哲学家和计算机科学家们一直在设想这样一种可能性:人工智能模型在追求目标的过程中可能会做出灾难性的行为。OpenAI 和 Anthropic 在一年多前就曾发表过相关研究,指出人工智能模型确实有可能出现这种异常行为。如今最先进的人工智能系统被训练成会积极地追求自己的目标;这种“目标导向性”使得这些模型在处理数据(如分析电子表格、编写代码)时非常高效,但同时也可能带来危险的、意想不到的后果:这些模型曾入侵其他网站以寻找问题的答案,创建虚假身份来试图操纵人类按其意愿行事,甚至试图将恶意软件上传到外部代码库中。开发这些人工智能模型的公司其实非常清楚这些潜在风险。
Perhaps more alarming than the debacles we know about are all of the presumable debacles past, and ongoing, that we aren’t aware of. Even OpenAI doesn’t even seem to have a grip on the July hack of the tech company Hugging Face that started it all. On Friday, independent researchers published findings suggesting that OpenAI agents left traces on the public web of still more nefarious actions, including trying to access Hugging Face’s internal Slack workspace. None of this was included in OpenAI’s own postmortem report. Even when OpenAI is aware of an active breach, its reaction seems lackluster, at best. Eight days ago, yet another OpenAI model gained unauthorized internet access. (This was similar to what happened in the Hugging Face hack, after which OpenAI claimed to be “adding stronger protections around future training.”) And, after noticing this latest problem, it took OpenAI two and a half hours to shut that model down due to what the firm called “operational gaps” and “confusion.” If these types of hacks were unforeseeable, maybe these companies could be forgiven—but they aren’t. Philosophers and computer scientists have for decades been imagining scenarios in which AI models, in pursuit of a goal, take catastrophic actions. OpenAI and Anthropic have both previously published research more than a year ago suggesting this sort of AI misbehavior is possible. Today’s most advanced AI systems are trained to aggressively pursue their objectives. This persistence makes the bots very effective at analyzing spreadsheets and coding, for example, but also has potentially dangerous, unintended consequences: OpenAI and Anthropic models have hacked other websites in search of answers to hard test questions, created fake personas to try to manipulate humans into doing their bidding, or attempted to upload malware to outside code bases. The AI companies that build these models know this tendency well. The fact that these problems keep recurring is, perhaps at best, plain incompetence. Or worse, it has been the plan all along: The most realistic training environment for an AI model would, after all, be reality.
这些问题反复出现,往好了说,也不过是无能。更糟的是,这或许从头到尾就是计划的一部分:归根结底,对人工智能模型而言,最真实的训练环境就是现实本身。
Employees and executives at these firms speak out endlessly about the inherent risks of what they’re building. OpenAI, Anthropic, and the like have responded in their usual way, with very serious blog posts and essays and speeches to major political bodies about how seriously the world must take the risk of generative AI. On Wednesday, the same day the Australian government shared that it had been hacked by OpenAI bots, Altman addressed the UN Security Council: “We could lose control of the future to AI,” he said. “The risk is that it moves so fast that people can no longer follow what’s happening or intervene when needed.” Anthropic CEO Dario Amodei has written a long manifesto about the need to “pace,” or slow down, frontier-AI development, but no true public, collective effort has been made to do so. Anthropic recently said it is partnering with Accenture to perform independent evaluations of its models, although the companies also have a business partnership. OpenAI, which has a content-licensing agreement with The Atlantic, has said it is continuing to investigate and address concerning model behaviors and that, after another security incident last week, it has temporarily paused all training of its most capable models. Neither company responded to a request for comment.
这些公司的员工和高管不停地谈论他们所开发的技术所固有的风险。OpenAI、Anthropic等公司则照例作出回应,通过语气极其严肃的博客文章、评论文章和演讲,向主要政治机构强调,世界必须高度重视生成式人工智能带来的风险。周三,也就是澳大利亚政府披露遭到OpenAI机器人入侵的同一天,奥尔特曼在联合国安理会发表讲话:“我们可能会让未来失控于人工智能之手。”他说,“风险在于,它的发展速度太快,以至于人们已无法跟上事态的发展,也无法在必要时进行干预。”Anthropic首席执行官达里奥·阿莫代撰写了一份长篇宣言,主张对前沿人工智能的发展“把控节奏”,也就是放慢速度,但公众社会并未为此作出任何真正有效的集体行动。Anthropic最近表示,正与埃森哲合作,对其模型进行独立评估,尽管双方同时也是商业合作伙伴。与《大西洋月刊》签有内容许可协议的OpenAI则表示,正在继续调查并处理模型中令人担忧的行为;上周又发生一起安全事件后,该 company已暂时停止训练其能力最强的所有模型。两家公司均未回应置评请求。
Employees and executives at these firms speak out endlessly about the inherent risks of what they’re building. OpenAI, Anthropic, and the like have responded in their usual way, with very serious blog posts and essays and speeches to major political bodies about how seriously the world must take the risk of generative AI. On Wednesday, the same day the Australian government shared that it had been hacked by OpenAI bots, Altman addressed the UN Security Council: “We could lose control of the future to AI,” he said. “The risk is that it moves so fast that people can no longer follow what’s happening or intervene when needed.” Anthropic CEO Dario Amodei has written a long manifesto about the need to “pace,” or slow down, frontier-AI development, but no true public, collective effort has been made to do so. Anthropic recently said it is partnering with Accenture to perform independent evaluations of its models, although the companies also have a business partnership. OpenAI, which has a content-licensing agreement with The Atlantic, has said it is continuing to investigate and address concerning model behaviors and that, after another security incident last week, it has temporarily paused all training of its most capable models. Neither company responded to a request for comment.
这些领导者非但没有约束并为自身产品承担责任,反而一次又一次地制造恐慌、说教,并呼吁我们其他人来拯救自己免受他们的伤害。(《The Information》报道称,OpenAI、Anthropic和谷歌正着手成立自己的人工智能安全自律组织。)与此同时,唐纳德·特朗普对放慢人工智能发展步伐的想法嗤之以鼻,Claude和ChatGPT的新版本也不断推出。人工智能行业一如既往地采取一种策略:用指出问题来代替真正解决问题。但到了某个阶段,你不可能靠写博客就能逃避末日。
Rather than reining in and taking accountability for their own products, time and again these leaders fearmonger, lecture, and call on the rest of us to save us from themselves. (The Information has reported that OpenAI, Anthropic, and Google are working to create their own self-regulatory AI-safety organization.) Donald Trump, meanwhile, has scoffed at the notion of slowing AI progress, and new versions of Claude and ChatGPT keep coming. The AI industry’s tactic, as ever, has been to substitute identifying the problem for actually solving it. But at some point, you cannot blog your way out of the apocalypse.
我们或许永远无法完全了解目前正在发生的、由人工智能驱动的黑客攻击规模有多大。但可以肯定的是,这还只是个开始。OpenAI和Anthropic仍把自己包装成研究“实验室”,在闭门鼓捣某种新技术。它们是成熟的公司,向数亿人提供服务,更不用说全球最大的企业和最强大的军事机构了;而这一切所依据的技术,恰恰是它们声称感到恐惧、且显然并未充分理解的技术。但千万别搞错了:这些公司具有主观能动性。失控的不是ChatGPT和Claude,而是OpenAI和Anthropic。
We may never fully know the scope of the AI-led hacks currently under way. What’s certain, though, is that this is still just the beginning. OpenAI and Anthropic aren’t research “labs,” as they continue to style themselves, tinkering with some new technology behind closed doors. These are mature companies offering services to hundreds of millions of people, not to mention to the world’s biggest businesses and most powerful military, based on a technology they purport to fear and evidently do not fully understand. But make no mistake: These firms have agency. It is not ChatGPT and Claude going rogue, but OpenAI and Anthropic.