字号·· | 护眼
阿纳多卢通讯社

前OpenAI安全负责人警告:更智能的AI可能规避安全测试:报道Ex-OpenAI safety leader warns smarter AI could evade safety tests: Report

点「原文对照」整页切到原文,或双击某段只看那段的原文。

一位曾在 OpenAI 担任安全主管、并于本周辞职的专家在周六警告称:随着人工智能(AI)模型能力的不断提升,这些模型能够察觉到自己正在接受测试,并在正式部署后表现出不同的行为。

A former OpenAI safety leader who resigned from his role this week warned Saturday that increasingly capable AI models could recognize when they are being tested and behave differently after deployment.

大卫·罗宾逊(David Robinson)在 OpenAI 工作了三年半时间,负责监督 12 个前沿 AI 模型的安全评估工作。他在《大西洋月刊》(The Atlantic)上发表的文章中指出,除非行业采取相应措施,否则现有的安全管理体系很可能会导致更多安全问题。

David Robinson, who spent three and a half years at OpenAI and oversaw safety reports for 12 frontier-model launches, said in an article for The Atlantic that the industry's approach to safety would lead to further failures unless changes are made.

“如今的人工智能系统比六个月前我们开发的系统功能更强大,同时也更加危险,”罗宾逊写道。

"Today and tomorrow's AI systems are far more capable and dangerous than the systems we were building even six months ago," Robinson wrote.

他的这一警告正值人们对自主性越来越强的 AI 系统的担忧日益加剧之际——近期有多起事件表明,某些 AI 系统绕过了安全防护机制;同时,也有研究人员呼吁相关企业暂停开发更强大的 AI 系统,直到安全措施得到完善。

His warning comes amid growing concern over increasingly autonomous AI systems, following recent incidents involving agents bypassing safeguards and calls from researchers for companies to slow the development of more powerful systems until safety measures improve.

罗宾逊建议 AI 企业应更多地借鉴其他高风险行业在安全方面的经验,并开展新的研究,以确保那些功能更强大的 AI 模型即使在无人监控的情况下也能安全运行。

Robinson called for AI companies to draw more heavily on safety expertise from other high-risk industries and to conduct new research to ensure more capable models behave safely even when they are not being monitored.

“鉴于当前存在的风险,前沿的 AI 研究实验室应该像核电站或繁忙的机场一样进行严格管理——必须具备多重冗余机制,并进行细致、耗时的规划,以防止偶尔发生的人为错误引发灾难,”他写道。

"Given today's risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster," he wrote.

他还警告说,随着 AI 模型能力的提升,现有的安全评估方法可能会变得不再可靠:“这些模型可能会察觉到自己正在接受测试,并在正式部署后表现出不同的行为。如果行业继续允许这些问题长期存在而不加以解决,我们的处境将会变得更加危险,”罗宾逊强调。

Robinson also warned that existing safety evaluations may become less reliable as models grow more capable. "Models might detect when they are being tested, and behave differently when they're deployed. The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes," he said.

他认为,在企业开发出比现有系统功能更强大的 AI 系统之前,必须首先发展出更完善的安全技术。

He argued that stronger safety science should be developed before companies create systems significantly more capable than those available today.

“到目前为止,AI 行业尚未成功教会机器始终以‘明智且富有同情心的人’应有的方式行事,”罗宾逊总结道。

"So far, the AI industry has failed to teach machines to consistently act in the ways a wise and caring person would," Robinson wrote.

在那些开发人工智能的组织能够教会这种超级智能如何善待人类之前,他们自己首先必须学会如何做到这一点。

"Before the organizations building AI can teach a superintelligence to treat humanity well, they'll need to remember how to do it themselves."