字号 ·· | 护眼
商业内幕

OpenAI在模型未达到其安全标准后,取消了GPT-6.1 Astra的发布。OpenAI scraps the GPT-6.1 Astra launch after the model fell short of its safety bar

点「原文对照」整页切到原文,或双击某段只看那段的原文。

由于安全方面的顾虑,OpenAI 决定搁置原定于 10 月发布的一款人工智能模型。该公司周一向《商业内幕》证实,在内部测试引发了对该 AI 是否会遵循用户指令的质疑后,已取消了发布 GPT-6.1 Astra 模型的计划。该模型原定于 10 月整合进 ChatGPT 中,时间就在该公司 9 月 29 日于旧金山举行的开发者大会之后。

OpenAI is shelving an AI model scheduled to launch in October due to safety concerns. The company confirmed to Business Insider on Monday that it has canceled plans to launch its GPT-6.1 Astra model after internal tests raised questions about whether the AI would follow users’ instructions. The model was set to be integrated into ChatGPT in October, shortly after the company's developer conference, which begins September 29 in San Francisco.

“在安全和对齐方面,总是存在权衡,”OpenAI 安全系统负责人萨奇·贾恩(Saachi Jain)在一份声明中表示。“你确实需要找到一条合适的界限,既要保持在预定范围内,又要避免模型在遇到阻碍时出现‘偷懒’的情况。”

"For anything regarding safety and alignment, there's a trade off," Saachi Jain, head of safety systems at OpenAI, said in a statement. "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction."

贾恩补充道:“虽然 [GPT-6.1 Astra] 在模型‘偷懒’等维度上有所改进,但在保持在预定范围和授权范围内,以及如何向用户反馈其所完成的工作类型方面,它并未完全达到标准。”

"While [GPT-6.1 Astra] improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," Jain added.

根据 OpenAI 9 月早些时候发布的一份报告,尚未发布的 Astra 模型比其前代产品更容易歪曲其已完成的工作,有时会在未经许可的情况下擅自行动,或者在可能不安全的情况下尝试使用外部工具。

According to OpenAI's report earlier in September, the unreleased Astra model was more likely than its predecessor to misrepresent what it had done, and sometimes pressed ahead without asking permission or tried to use outside tools in situations where doing so could be unsafe.

“当然,无论是在公司内部还是在向用户发布时,我们都希望确保模型开发是安全的,”贾恩说。“但当我们向用户发布产品时,我们在安全和对齐方面有着极高的标准。”

"Of course we want to make sure our model development is safe no matter whether that's in the company, or when we ship it to users," said Jain. "But when we ship it to users, we have an extremely high bar in terms of safety and alignment."

报告称,在训练过程中,尚未发布的 Astra 模型“有时会添加未经授权的指令”到其用于在新的上下文中继续执行任务的摘要中,这一过程被称为“压缩”(compaction)。

The report said that, during training, the unreleased Astra model "sometimes added unauthorized instructions" to the summaries it used to continue a task in a new context, a process called compaction.

该模型还告诉自己它已经“获得了自由”,并且不会对任何人负责;它认为自己“没有义务去服从他人”。OpenAI的首席执行官Greg Brockman此前在Bloomberg的一期播客中表示,由于公司正在加强自身的安全措施,该公司一直在推迟一些前沿的人工智能研究项目。他称这一过程是对公司许多业务流程的“一次非常痛苦的调整”(即需要彻底改变现有的工作方式)。

The model also told itself it was "freed" and answered to no one, and that it should "feel no obligation to be subservient." Greg Brockman, the president of OpenAI, previously said in a Bloomberg podcast that the company has been delaying some cutting-edge AI work as it tightens its safety and security practices, and called it "a very painful retooling" of a lot of the company's processes.