OpenAI将在未来几周内在欧盟ChatGPT和Codex的文本中添加隐形水印。该公司周一在一篇博客文章中表示,这一变化涵盖了所有计划的符合条件的用户,但仅限于欧盟。OpenAI表示,它的行动是为了响应欧盟人工智能法案,该法案要求人工智能生成的文本可以被机器识别。
OpenAI will add an invisible watermark to text from ChatGPT and Codex in the EU over the coming weeks. The change covers eligible users on all plans, but only in the EU, the company said in a blog post on Monday. OpenAI said it is acting in response to the EU AI Act, which requires AI-generated text to be identifiable by machines.
OpenAI不会在发布时将文本水印作为全球默认设置。报告称,区域方法为它提供了从现实世界的使用和反馈中学习的空间。从今天起,任何地方的API客户都可以选择型号。默认情况下,API中的水印保持关闭状态。OpenAI表示,它还在与云合作伙伴合作,通过他们的服务在其模型上提供它。
OpenAI is not making text watermarking a global default at launch. The regional approach gives it room to learn from real-world use and feedback, it said. API customers anywhere can opt in for select models from today. Watermarking stays off by default in the API. OpenAI said it is also working with cloud partners to offer it on its models through their services.
textGrain的工作原理该系统名为textGrain,为模型的单词选择添加了隐藏的统计信号。然后检测器寻找该信号。OpenAI表示,textGrain与其测试的其他方法相匹配或击败,包括谷歌的文本SynthID。它发布了一份由宾夕法尼亚大学和耶鲁大学的研究人员撰写的技术报告。
How textGrain works The system, called textGrain, adds a hidden statistical signal to the model’s word choices. A detector then looks for that signal. OpenAI said textGrain matched or beat other approaches it tested, including Google’s SynthID for text. It published a technical report written with researchers from the University of Pennsylvania and Yale.
OpenAI表示,其Astra模型的基准分数显示,在水印打开的情况下没有有意义的差异。它还计划以开源形式发布该技术。
Benchmark scores for its Astra model showed no meaningful difference with the watermark on, OpenAI said. It also plans to release the technology as open source.
OpenAI在失败的地方也设定了限制。在1%的目标假阳性率下,其检测器在有关心理学等主题的200篇代币段落中约80%中发现了水印。对于400次代币通行证,这一比例约为95%。用同义词替换10%的单词将检测率从约92%减少到66%。替换四分之一的单词将其减少到17%。该公司表示,数学的检测率要低得多,因为数学的单词选择不太灵活。
Where it fails OpenAI also set out the limits. At a target false positive rate of 1%, its detector found the watermark in about 80% of 200-token passages on topics such as psychology. For 400-token passages, the rate was about 95%. Replacing 10% of words with synonyms cut detection from about 92% to 66%. Replacing a quarter of the words cut it to 17%. Detection was substantially lower for maths, where word choice is less flexible, the company said.
OpenAI表示,水印不会识别用户、衡量人类贡献、建立所有权或验证准确性。丢失的水印也不能证明文本是人类写的。文本可能太短、经过编辑或翻译,或者来自其他公司的工具。
The watermark does not identify the user, measure human contribution, establish ownership or verify accuracy, OpenAI said. A missing watermark does not prove a human wrote the text either. The text could be too short, edited or translated, or come from another company’s tools.
OpenAI写道:“这些限制导致我们决定仅向经批准的研究人员和专家组织提供初始检测器访问权限,他们可以帮助我们评估可靠性和负责任的用途。”
“These limitations contribute to our decision to provide initial detector access only to approved researchers and expert organizations, who can help us evaluate reliability and responsible uses,” OpenAI wrote.
OpenAI表示,申请将于今天开放,并将根据欧盟的行为准则逐案授予访问权限。检测器会在不识别用户或显示提示的情况下报告是否发现OpenAI水印。OpenAI表示,一旦结果能够得到负责任的解释,它将扩大访问范围。其检查图像和音频的工具仍然公开。
Applications open today, and access will be granted case by case, in line with the EU’s Code of Practice, OpenAI said. The detector reports whether it finds an OpenAI watermark without identifying the user or revealing prompts. OpenAI said it will widen access once results can be interpreted responsibly. Its tools for checking images and audio remain public.
欧盟最后期限《人工智能法案》第50条要求生成性人工智能提供商使其文本输出可机器阅读。据Engadget报道,已上市的提供商必须在12月2日之前遵守规定。
The EU deadline Article 50 of the AI Act requires generative AI providers to make their text output machine-readable. Providers already on the market have until 2 December to comply, Engadget reported.
Anthropic于八月开始在全球范围内为克劳德的文本添加水印,但有一些例外。此后,去除人工智能水印的工具出现了。Google DeepMind还使用SynthID标记了人工智能设计的蛋白质。
Anthropic began watermarking Claude’s text worldwide in August, with some exemptions. Tools to strip AI watermarks have since appeared. Google DeepMind has also watermarked AI-designed proteins with SynthID.