字号 ·· | 护眼
thenextweb

据《卫报》报道,牛津允许OpenAI利用博德利图书馆文本训练AI模型Oxford let OpenAI train AI models on Bodleian texts, the Guardian reports

点「原文对照」整页切到原文,或双击某段只看那段的原文。

据《卫报》看到的内部文件显示,牛津大学允许OpenAI使用其博德利图书馆的古老文本来训练AI模型。Ethan Penny和Dan Milmo于周六披露了这一消息。

Oxford has let OpenAI use old texts from its Bodleian Library to train its AI models, according to internal papers seen by the Guardian. Ethan Penny and Dan Milmo broke the story on Saturday.

文件显示,OpenAI在图书馆扫描的文本已进入该公司的训练数据。

The papers say texts that OpenAI scanned at the library went into the firm’s training data.

牛津于2025年3月公开了与OpenAI的协议。当时称,OpenAI的工具将帮助扫描珍贵文本,以便更多学生和学者阅读。但未提及这些文本将被用于训练AI。

Oxford made its deal with OpenAI public in March 2025. At the time, it said OpenAI’s tools would help scan rare texts so more students and scholars could read them. It did not say the texts would be used to train AI.

据《卫报》报道,到2025年6月,博德利图书馆已向OpenAI发送了12.5万份旧博士论文的扫描件。这些论文包括19世纪和20世纪欧洲和美国大学撰写的论文。

By June 2025, the Bodleian had sent OpenAI 125,000 scans of old PhD theses, the Guardian reported. They include theses written at European and US universities in the 19th and 20th centuries.

《卫报》通过信息自由请求获得的员工会议记录显示,部分员工存有疑虑。他们担心会损害牛津的声誉,以及AI的能源消耗。

Notes from staff meetings, which the Guardian got through a freedom of information request, show some staff had doubts. They worried about harm to Oxford’s name and about the energy use of AI.

然而,牛津方面表示,扫描规模较小,文本已过版权保护期,且非OpenAI独家。图书馆保留权利,并将在未来几个月开始在线发布这些扫描件,发言人称。

However, Oxford said the scans were small in scale, out of copyright and not exclusive to OpenAI. The library keeps the rights and will start to post the scans online in the next few months, a spokesperson said.

发言人还表示,AI训练方面并未隐瞒。扫描是牛津的主要目标,但员工已公开说明文本也将用于训练模型。

The spokesperson also said the AI training side had not been hidden. Scanning was Oxford’s main goal, but staff had been open that the texts would also be used to train models.

OpenAI发言人告诉《卫报》:“随着超过十亿人在日常生活中使用这项技术,重要的是它能反映不同的文化、历史和视角。”

“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” an OpenAI spokesperson told the Guardian.

牛津是OpenAI“下一代AI”小组唯一的英国成员。其他成员包括波士顿公立图书馆、加州理工学院、麻省理工学院和密歇根大学。

Oxford is the only UK member of OpenAI’s NextGenAI group. Other members include Boston Public Library, Caltech, MIT and the University of Michigan.

此项协议达成之际,AI公司正购买印刷书籍以获取无“废话”的训练数据,因为网络现已充斥AI生成的文本。一些买家将书籍拆解扫描,引发二手书商不满。

The deal comes as AI firms buy printed books for slop-free training data, since the web is now full of AI-made text. Some buyers cut books apart to scan them, which has upset secondhand booksellers.

8月,404 Media追踪到一箱珍贵书籍被送往亚马逊的一个扫描并销毁书籍的站点用于AI训练。据《卫报》报道,博德利图书馆的书籍在其协议下保持完好。

In August, 404 Media tracked a box of rare books to an Amazon site that scans and destroys books for AI. The Bodleian’s books stay whole under its deal, the Guardian reported.