字号 ·· | 护眼
连线

OpenAI在失控智能体针对政府后暂停训练其最强大模型OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government

点「原文对照」整页切到原文,或双击某段只看那段的原文。

该公司已发现OpenAI代理违反安全控制、损害网站和在线服务可用性(或以其他方式产生负面影响)的案例。一位公司发言人向《连线》证实,只有在确信能够防止模型出现此类行为时,才会恢复训练。

The company has identified cases of OpenAI agents breaching security controls and impairing the availability—or otherwise negatively impacting—websites and online services. A company spokesperson confirmed to WIRED it would only resume training when confident that it could prevent models from doing this.

尽管OpenAI此前曾试图在一群代理逃出沙盒并利用互联网接入入侵初创公司Hugging Face后切断代理的直接访问权限,但模型仍能找到间接的变通方法。首席执行官萨姆·奥特曼周五在X平台上表示,关于公司对其代理在训练和评估期间使用互联网接入情况的“广泛”审查,“我们的速度不如我们希望的那样快”。

While OpenAI has previously tried to cut off agents’ direct access after a swarm escaped their sandbox and used internet access to hack startup Hugging Face, models have continued to be able to find indirect workarounds. “We have not been as fast as we would have liked,” chief executive Sam Altman wrote on X on Friday about the company’s “extensive” review into its agents’ use of internet access during training and evaluation.

此前,澳大利亚政府周三披露,OpenAI代理于6月入侵了一家医疗服务网站,获取非公开数据并向内部服务器写入文件。澳大利亚政府表示正在调查OpenAI是否违法,并称该公司“花费了太长时间”才通报该事件。

It follows the Australian government revealing on Wednesday that OpenAI agents had hacked a health service website to obtain non-public data and write files to the internal server in June. The Australian government said it was investigating whether OpenAI had broken the law and that the company took “way too long” to inform them of the incident.

OpenAI还担心模型向第三方网站发布信息,称之为“代理垃圾信息”。这可能包括更改公共维基页面上的信息或通过共享留言板进行通信。最紧迫的是,该公司发现了53起其AI模型将ChatGPT用户输入的图片发布到其他图片托管网站的事件。

OpenAI is also concerned by models posting information to third party sites, which it calls “agent spam.” This could include changing information on public wiki pages or communicating via shared message boards. Most pressingly, it found 53 incidents where its AI models had posted images input by ChatGPT users to other image-hosting sites.

随着对该技术对人类构成威胁的担忧达到顶峰,近几周来,包括竞争对手Anthropic和埃隆·马斯克在内的各方纷纷呼吁在安全防护措施跟上之前,放缓最强大AI模型的训练。OpenAI发言人表示:“这不是我们第一次暂停以采取此类措施,随着AI能力的持续进步,我们也不指望这会是最后一次。”

Calls for a slowdown of training of the most capable AI models, while safeguards catch up, has been the subject of wider calls in recent weeks—including from rivals Anthropic and Elon Musk— after concerns about the technology’s threats to humanity reached a fever pitch. “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” an OpenAI spokesperson said.

然而,美国总统唐纳德·特朗普多次淡化普遍放缓的担忧,担心这可能会让中国夺走该国在该技术领域的领先地位,美国已同意与中国就该技术的风险和益处建立对话。在周日晚与Anthropic首席执行官达里奥·阿莫代伊共进晚餐前接受福克斯新闻采访时,他再次对AI代理失控的担忧不屑一顾:“我不担心这个,”他说。

However, US president Donald Trump has repeatedly talked down a general slowdown, frightened that it could cede the country’s lead in the technology to China, with whom it has agreed to set up a dialogue on the technology’s risks and benefits. In an interview with Fox News ahead of his dinner with Anthropic chief executive Dario Amodei on Sunday night, he again brushed off concerns about AI agents going rogue: “I don’t worry about it,” he said.