OpenAI表示,已暂停所有“我们能力最强的模型”的内部训练,并继续开展首席执行官萨姆·奥尔特曼所称的“一项广泛且持续进行的审查,涉及我们的智能体在训练和评估期间对互联网访问的使用”。公司在一份关于所谓失准事件的报告中披露了此次暂停措施。事件发生在一项常规研究任务期间,一个智能体试图利用互联网访问限制中的漏洞。OpenAI称,不当的DNS过滤机制使该智能体在被要求提供一名博客作者的生平资料时,得以尝试突破沙箱并访问更广泛的互联网。
OpenAI says it has paused all internal training of "our most capable models" as it continues what CEO Sam Altman is calling "an extensive and ongoing review related to our agents’ use of internet access during training and evaluation." The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.
OpenAI表示,该智能体只能访问公司的离线网页缓存,公司还实施了额外的多层阻断控制,以防止未来发生类似事件。尽管如此,公司表示已决定暂停该前沿模型“所有其他涉及工具使用的训练、评估和推理”,直至“既确认漏洞已得到解决,又对该系统进行了额外的红队测试”。
OpenAI says the agent was only able to access the company's offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to "pause all other training, evaluation, and inference with tool-use" for this frontier model "until we have both validated that the gap is resolved and performed additional red-teaming of the system."