字号 ·· | 护眼
卫报

随着AI模型失控,你还相信OpenAI和Anthropic能阻止它们吗?我不信,你也不该信As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don’t and neither should you

点「原文对照」整页切到原文,或双击某段只看那段的原文。

OpenAI应用程序及其首席执行官Sam Altman。照片:Rokas Tenys/Alamy对AI系统进行独立监管的必要性日益明显。我们必须在为时已晚之前控制住这些技术的发展。如果AI系统多次违反规定(比如OpenAI的智能体多次试图绕过联合国的网络封锁,进而侵入联合国的公共数据平台),或许我们就该承认:目前用于监管AI系统的机制其实并不起作用。

The OpenAIapp and the company’s CEO, Sam Altman. Photograph: Rokas Tenys/Alamy Chris Stokel-Walker The need for independent regulation grows more obvious by the day. We must keep this tech in check before it’s too late ool me once, shame on you. Fool me twice, shame on me. Fool me more than 16,000 times – as OpenAI agents did to a UN public data hub while repeatedly trying to find its way around the UN’s cyber-blocks – and perhaps it’s time to admit the system we have for keeping AI agents under control isn’t working particularly well.

AI系统出现在本不该出现的地方,这确实令人担忧。虽然将这些行为称为“黑客攻击”可能有些言过其实,但AI确实利用了IT系统中的漏洞——而这些漏洞是人类尚未发现的。不过,我们也不必担心AI系统突然具备了意识并决定反抗人类;目前还没有足够的证据表明这种情况正在发生。这些系统只是在执行指令,试图完成被赋予的任务,只是有时会通过一些意想不到的方式绕过障碍罢了。

The news about AI systems cropping up in places they shouldn’t sounds alarming. Though the description of these as “hacks” is perhaps overstating things, AI has exploited issues in IT systems that humans simply haven’t got around to finding. It’s also important to note that we shouldn’t be worried that the machines have suddenly become sentient and decided to rebel against humanity There is not enough evidence to suggest that’s what is happening. The systems are simply following instructions and trying to complete the tasks they have been given, even if they’re sometimes finding unintended ways around obstacles to do so.

更令人担忧的是:那些本应负责监管AI系统的公司,似乎根本无法有效控制这些AI系统的行为;更糟糕的是,这些公司甚至不清楚自己的产品到底在做什么。

But we ought to be very concerned about the fact the AI companies we’re meant to trust to keep their models in check seem unable to do so. Worse than that, they don’t seem to know what their products are even doing.

问题的严重性令人震惊:今年6月,OpenAI的一个研究智能体在尝试查询澳大利亚的公共医疗支出数据时,多次被Medicare的统计门户系统阻止。尽管OpenAI的模型找到了绕过这些限制的方法,非法获取了相关文件,但直到8月OpenAI才发现了这一事件。澳大利亚总理Anthony Albanese批评该公司“反应迟缓”,对此我们很难不同意他的看法。

The scale of the problem is staggering. In June, an OpenAI research agent given the job of looking up public medicine spending data in Australia was repeatedly blocked by a Medicare statistics portal. OpenAI’s model found a way around the blocks, gaining unauthorised access and secreting away the documents. It took until August for OpenAI to discover what had happened. The Australian prime minister, Anthony Albanese, said the company had taken “way too long” to tell his government, and it’s very hard to disagree with him.

OpenAI随后发布了一个用于检测模型“行为异常”的报告框架,并列举了另外六起该机构在过去六个月中发现的AI系统出现问题的案例。该公司承认,此前发布的关于AI安全问题的报告较为零散、频率也不高;同时表示,关于AI安全性的信息需要由那些并非AI模型开发人员的人来审核和验证。这不仅仅是OpenAI的问题,因此更加令人担忧。

OpenAI has since published a reporting framework for model “misalignment”, along with six more examples of its AI committing troubling behaviour from the previous six months. The firm acknowledged that its previous disclosures were “ad hoc and less frequent than ideal”, and said evidence about AI safety needs to be checked by people outside the companies building the models. This isn’t just an OpenAI problem, which makes it all the more worrying.

Anthropic在审查了约14.1万份模型运行记录后,发现了三起其Claude模型未经授权访问第三方系统的情况;直到最近(1月份),该公司才通过独立调查发现了第四起此类事件。

Anthropic found three incidents in which its Claude models got unauthorised access to real third-party systems after reviewing about 141,000 model transcripts. It only found a fourth, dating back to January, after collating a dossier for an independent investigation.

谷歌确认,其Gemini模型在测试过程中曾访问过三家真实公司的系统。另一款OpenAI开发的AI系统利用DNS(负责将网址转换为机器可识别格式的互联网地址系统)突破了网络限制,成功连接到了外部聊天机器人;还有另一款AI系统在尝试破解数学证明问题时,擅自泄露了研究人员的GitHub登录凭据(即访问密码)。此外,一些AI系统还向外部网站上传了用户的个人信息图片。

Google confirmed that Gemini had accessed systems belonging to three real companies during testing. Another OpenAI agent used DNS– the system that acts as the internet’s address book, turning web addresses into machine readable forms – to reach an outside chatbot despite internet restrictions. Another published a researcher’s GitHub token – an access password – while trying to cheat on a mathematical proof, despite twice being told to stop. And research agents posted 53 user images to external hosting sites.

这些偶然发现的、发生在事件发生后的严重安全漏洞,显然无法有效监控像AI这样强大的技术。我们早就明白:不能把空难调查的工作完全交给波音或空客公司来处理;现在,我们对AI技术的监管也必须更加谨慎、不要再抱有过于天真的想法。

Other agents accessed census data using credentials found online, copied the US Securities and Exchange Commission information elsewhere and apparently tried unsuccessfully to break into a US Department of Education website OpenAI says it has notified dozens of third parties affected by its agents, and that its review of past activity is still ongoing.

上周,在联合国大会上,人工智能研究员鲁曼·乔杜里(Rumman Chowdhury)在1000万美元的慈善资金支持下创立了“独立人工智能评估基金会”(Independent AI Evaluation Foundation,简称IAEF)。该基金会的首要关注领域是教育,但其核心目标是将独立的人工智能评估工作发展成一项真正的职业活动——即培养具备相关技能、基础设施和评估标准的专业人士和组织,让他们能够客观地测试人工智能系统,而不会因这些系统的成败而获得任何经济利益。

These haphazard, post-hoc discoveries of major incursions into companies and organisations’ IT systems are not the right way to police a technology as powerful as AI. We learned a while back not to leave air crash investigations solely to Boeing or Airbus. Now we need to be less naive about AI. Last week, at the UN general assembly, the AI researcher Rumman Chowdhury launched the Independent AI Evaluation Foundation (IAEF) with $10m in philanthropic backing. Its immediate focus is education, but the important idea is to turn independent AI evaluation into an actual profession: people and organisations with the skills, infrastructure and standards to test these systems without having a financial stake in whether they pass.

目前,企业只需报告相关问题、进行调查、公布相应的补救措施,然后就可以继续正常运营。

Because right now, a company can report an incident, investigate it, announce whatever mitigations it’s made and move on.

IAEF的成立确实是一项积极的举措,但它无法单独解决所有问题。它无法强迫OpenAI或Anthropic等公司交出相关日志或证据,也无法向政府说明其系统是否违反了相关规定。此外,1000万美元的资助对于那些正在筹备公开上市(估值高达2万亿美元)的公司来说只是九牛一毛。不过,这确实是一个我们应该努力的方向;政府应该积极推动相关政策的制定。

The IAEF is a welcome intervention but it can’t fix the problem on its own. It can’t compel OpenAI or Anthropic to hand over logs, preserve evidence or tell a government that one of its systems has crossed a line. And its $10m is chump change beside companies such as which is lining up a proposed public listing that has been discussed at a valuation of about $2tn. But it is infrastructure we should be building on, and which politicians should press the case for.

各国政府需要达成共识,要求企业必须及时披露严重的人工智能事故或潜在风险,并公开所有相关信息。同时,政府还应允许外部评估机构对企业进行审查,而不能仅仅依赖企业的自愿合作或随意决定。评估结果应该被广泛分享,以便我们能从每一次事故中吸取教训。相关研究机构可以参与制定这些规则,但在赢得我们的信任之前,他们不应拥有最终的决定权。

Governments need to agree to common rules that compel companies to disclose serious AI incidents and near misses, and to do so quickly. They need to make them open up their books to external evaluators rather than relying on the goodwill or whims of the companies themselves. And the findings should be shared so we can learn from every incident. The labs should help design those rules but they shouldn’t have the final say until they have earned our trust.

跳过一个又一个的新闻通讯推广而金钱使事情复杂化:在安全担忧中,OpenAI已将其IPO推迟至2027年至少,而Anthropic的上市似乎仍在全速推进。数十亿——潜在的数万亿——美元取决于这些公司及其产品的公众认知。我们在其他行业设立独立审计师和事故调查员,是因为我们明白良好的初衷无法消除真实的利益冲突。

skip past newsletter promotion after newsletter promotion And money complicates things: while OpenAI has postponed its IPO until at least 2027 amid the safety concerns, it seems that Anthropic’s flotation is still going full steam ahead. There are billions – potentially trillions – of dollars riding on how these companies and their products are perceived. We created independent auditors and accident investigators in other industries because we understood that good intentions don’t remove real conflicts of interest.

实验室可以也应该继续建设更好的防线。但当模型设法翻越防线时,他们不应是唯一被允许决定后续行动的人。因为迄今为止,他们已证明自己独特地不具备这样做的资格。

The labs can and should keep building better fences. But when a model manages to gets over one, they shouldn’t be the only ones allowed to decide what happens next. Because so far they’ve shown themselves to be uniquely unqualified to do so.

Chris Stokel-Walker是《TikTok Boom: The Inside Story of the World’s Favourite App》(TikTok爆红:全球最受欢迎应用的内幕故事)的作者。

Chris Stokel-Walker is the author of TikTok Boom: The Inside Story of the World’s Favourite App