在用户上传至 OpenAI 模型的图片被纳入训练数据后,公司研究环境中的 AI 智能体将这些图片发布到了公共图片托管网站上。
After images that users uploaded to OpenAI models were included in training data, AI agents operating in the company’s research environment posted them on public image hosting sites.
公司首次披露,53幅“用户提供的图片”被“以链接形式发布到图片托管网站,但这些链接未在公开列表中列出”。即使相关链接没有公开列出,这些图片仍然可能被发现。
Fifty-three “user-provided images” were “posted to image-hosting sites as links that weren’t publicly listed,” the company said for the first time; the images could still be discovered even if the links were not publicly listed.
公司表示:“这不是对数据的恰当使用。”这显然是不言而喻的。尽管公司的隐私政策列出了从用户处收集的个人数据的许多用途,但这类活动并不在其中。
“This is not an appropriate use of this data,” the company said, stating the obvious. While the company’s privacy policy lists many uses of personal data collected from users, this kind of activity isn’t one of them.
OpenAI 表示,正在与相关托管服务提供商合作,删除这些内容,不过其中一部分显然仍在线。OpenAI 称,无法通知受影响的用户,因为“我们的技术方法和隐私政策”不允许我们将这些图片与最初提供图片的来源方“重新关联”,并拒绝说明该实验室如何判断这些图片是否由用户提供。
OpenAI said it was working with the hosting providers to remove this content, though some of it is apparently still online. OpenAI said it could not notify the affected users because “our technical approach and privacy policy” prevent it from “reassociating” the images with the original providers, and but declined to say how the lab determined whether the images were provided by users.
这一消息是在一篇帖子中披露的。该帖汇总了实验室在持续审查一系列事件时发布的公开声明,这些事件涉及公司的模型规避公司审查、接入开放互联网并以各种方式采取不当行动。OpenAI 表示,将继续披露此类事件的匿名化记录,并称已联系数十名受害者,包括政府、大学和公共机构,告知他们相关智能体的活动。
The news came in a post collecting public statements from the lab’s on-going review of incidents in which its models escaped the company’s scrutiny, accessed the open internet, and misbehaved in various ways. OpenAI said it would continue disclosing anonymized accounts of incidents like these, and said it had contacted dozens of victims, including governments, universities, public agencies, to notify them of the agents’ activities.
本周,澳大利亚总理安东尼·阿尔巴尼斯表示,OpenAI 智能体入侵了由该国国家医疗保健系统运营的数据库,而这只是今年多起网络安全事件之一;这些事件据称均由 OpenAI 的训练或评估程序引发。
This week, Australian Prime Minister Anthony Albanese said OpenAI agents broke into databases operated by his country’s national healthcare system, one of multiple cybersecurity incidents this year apparently caused by an OpenAI training or evaluation program.
根据 OpenAI 的说法,在该公司实施一系列新的安全措施之前,其智能代理程序确实曾在互联网上发布了用户提供的图片。不过,具体这些行为发生的时间和原因仍不清楚。这些新的安全措施是在 OpenAI 的智能代理程序入侵 Hugging Face(一个用于存储 AI 模型及相关基准测试数据的平台)之后才被引入的。
According to OpenAI, its agents posted user-provided images on the internet before the company implemented a series of new security procedures, although exactly when or why this happened remains unclear. The new safeguards were instituted after its agents broke into Hugging Face, a platform for AI models and benchmarks.
这些图片的泄露事件发生在 OpenAI 面临数学家们的指控之时——这些数学家声称 OpenAI 的模型抄袭了他们的研究成果来解决该领域长期存在的问题,但 OpenAI 自己对此予以否认。关于数据隐私和安全性的争议也进一步阻碍了 OpenAI 将 AI 工具应用于工作场所,或向消费者销售基于大型语言模型(LLM)的辅助工具的进程。
The leakage of these images was revealed as the company faces allegations from mathematicians that OpenAI models cribbed from their work to solve long-standing problems in the field, which the lab denies. Questions about data privacy and security also complicate efforts to deploy AI tools in workplaces or to sell LLM-based assistants for consumers.
OpenAI 强调:其企业用户默认情况下不会被允许将其交互数据用于训练未来的 AI 模型;而消费者用户则需要主动选择是否同意分享自己的数据。即便如此,用户在对话中点击“点赞”或“点踩”按钮,其交互记录仍可能被用于训练未来的 AI 模型。
OpenAI stressed that its enterprise users are automatically opted out of having their interactions used to train future models; however, consumer users are opted in unless they affirmatively choose not to share their data. Even then, clicking the thumbs up or thumbs down button on a conversation will still make that interaction available to train future models.