字号 ·· | 护眼
theregister

大AI的内容难题:拿走作品,留下收益Big AI's content problem: Take the work, keep the money

点「原文对照」整页切到原文,或双击某段只看那段的原文。

大型AI公司面临一个问题。它们必须吞噬他人的作品。你知道的,书籍、新闻报道、照片、代码、网站、你想出的那个原创梗图,以及互联网上几乎所有其他内容,都被用来训练其大型语言模型(LLM)。我们都知道这一点。我们也知道,AI公司极其厌恶为此付费、遵守附带的许可协议,更别提与最初创作这些作品的公司和个人分享收益了。最近,大型AI的默认经营模式——“先拿走,再争辩合法性”——变得比以往任何时候都更加赤裸裸。例如,看看《纽约时报》与OpenAI和微软之间的版权之争。根据404 Media对近期解封法庭文件的报道——原告的论点,法院尚未对此作出裁决——微软据称在“大肆导入”互联网内容时明知故犯。微软应用科学总监Brent Hecht博士在原告提交的92页合并简报中被引述称:此案关乎“规模空前的惊人盗窃”。他在记者律师提交的简易判决简报[PDF]中被引述称,这甚至可能是“人类历史上最大规模的劳动成果盗窃”。请注意,这不是媒体人士的评论,而是微软高级员工的发言。Hecht并非微软唯一发表评论的人。在法庭文件中引用的一份微软内部政策文件里,作者承认生成式AI可能会“显著扰乱那些生成基础模型训练数据的人的就业[因为]LLM是一种摧毁其供应链的产品。”这进而导致模型崩溃。具有讽刺意味的是,AI正在扼杀那只下出有价值内容金蛋的内容创作者鹅。Sidney H.法官美国纽约南区联邦地区法院法官斯坦因尚未就该案作出裁决。OpenAI 和微软正以“合理使用”抗辩为由,积极反驳原告的指控。他们主张,利用公开的文章和书籍训练大语言模型有助于推动公共知识进步,并不构成“非法的经济市场替代品”。但文件中引用的言论是真实的。大型 AI 公司清楚自己需要这些内容。一些相关公司根本不在乎;他们的座右铭是“先赚钱,后担忧”。那些能看透营收数字的人非常清楚,他们正陷入微软在解密法庭文件中自称的 AI 内容战略“厄运循环”。这会阻止他们吗?不会。

Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that. We also know that AI companies hate paying for any of it, obeying the licenses attached to it, or, God forbid, sharing their revenue with the companies and people who created the work in the first place. Recently, Big AI's default way of doing business: "Take it now, argue about legality later," has become more in your face than ever. Look, for example, at the copyright fight between The New York Times and OpenAI and Microsoft. According to 404 Media’s reporting on recently unsealed court documents - arguments from the plaintiffs that the court has not yet ruled on - Microsoft allegedly knew what it was doing when it was importing the internet willy-nilly. Microsoft's Director of Applied Science, Dr. Brent Hecht, was quoted in the news plaintiffs' 92-page combined brief as saying: the case was about “an astonishing theft of unprecedented proportions." Indeed, it was possibly the “largest theft of labor in human history,” he was quoted as saying in the summary judgment brief from the journalists' lawyers [PDF].This, mind you, wasn't a comment by a member of the press; it was from a senior Microsoft staffer. Hecht wasn't the only one at Microsoft who commented. In an internal Microsoft policy document also quoted in the court papers, the authors admitted generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain.” That, in turn, leads to model collapse. Ironically, AI is killing the content-creator goose that lays the gold eggs of worthwhile content. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not yet issued a decision on the case. OpenAI and Microsoft are actively disputing the plaintiffs' allegations under a "fair use" defense. They contend that using public articles and books to train large language models helps push forward public knowledge, and isn't acting as a "unlawful economic market substitute". But the statements the documents cites are real. Big AI knows it needs the content. Some of the companies concerned don't give a damn; money now, worry later is their motto. Those who can look past the revenue numbers are well aware they're engaged in what Microsoft itself called a “doom loop” of AI content strategy in those unsealed court documents. Will that stop them? Nah. According to the filing [PDF], which refers to sworn deposition testimony of OpenAI’s own corporate representative, it also testified that its LLMs were happy to vacuum up other people's work even when a paywall nominally protected it. Its rep was quoted as saying they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When OpenAI cofounder Greg Brockman was told OpenAI could hack its way through the firewall, he responded, “ah, nice.” Whether or not these statements are found to be reflective of a wider attitude, similar attitudes prevail amongst Big AI. Such policies could eventually come back and bite the LLM makers; it's already affecting the market for journalism, fiction writing, graphics creation, anything that demands humans actually get paid for producing original work. But Big AI just doesn't want to pay for it. That's not a side effect. It's the business model. Joe User couldn't care less. They just want a quick answer that sounds right. Some of them couldn't care less about getting the right answers. As for looking deeper to see if the information they're regurgitating is accurate, forget about it! As an OpenAI software engineer put it: “No matter how prominently we show the links, users won’t click.” Well, I click. That's a big reason why Perplexity is my AI of choice. It's not that it gives a more reliable answer. No, it's that, unlike most LLMs, Perplexity provides sources and links, and I check them before accepting what it tells me. But then I'm a journalist whose degrees are in history, where my professors drummed into me that you always - always - look at the primary sources. Big AI takes the same approach with its code generators. The US Ninth Circuit recently handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub lawsuit. The court ruled that the plaintiffs hadn't established a claim under one particular provision of the Digital Millennium

根据起诉文件 [PDF],文件引用了 OpenAI 企业代表的宣誓证词,证词称其大语言模型乐于吸纳他人作品,即使这些作品名义上受付费墙保护。该代表被引述称,不知晓“任何在训练数据集中检测付费墙内容的努力”,也不知晓“任何从训练数据集中移除付费墙内容的努力”。当 OpenAI 联合创始人格雷格·布罗克曼得知 OpenAI 可以黑客手段绕过防火墙时,他回应道:“啊,好极了。”

Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that. We also know that AI companies hate paying for any of it, obeying the licenses attached to it, or, God forbid, sharing their revenue with the companies and people who created the work in the first place. Recently, Big AI's default way of doing business: "Take it now, argue about legality later," has become more in your face than ever. Look, for example, at the copyright fight between The New York Times and OpenAI and Microsoft. According to 404 Media’s reporting on recently unsealed court documents - arguments from the plaintiffs that the court has not yet ruled on - Microsoft allegedly knew what it was doing when it was importing the internet willy-nilly. Microsoft's Director of Applied Science, Dr. Brent Hecht, was quoted in the news plaintiffs' 92-page combined brief as saying: the case was about “an astonishing theft of unprecedented proportions." Indeed, it was possibly the “largest theft of labor in human history,” he was quoted as saying in the summary judgment brief from the journalists' lawyers [PDF].This, mind you, wasn't a comment by a member of the press; it was from a senior Microsoft staffer. Hecht wasn't the only one at Microsoft who commented. In an internal Microsoft policy document also quoted in the court papers, the authors admitted generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain.” That, in turn, leads to model collapse. Ironically, AI is killing the content-creator goose that lays the gold eggs of worthwhile content. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not yet issued a decision on the case. OpenAI and Microsoft are actively disputing the plaintiffs' allegations under a "fair use" defense. They contend that using public articles and books to train large language models helps push forward public knowledge, and isn't acting as a "unlawful economic market substitute". But the statements the documents cites are real. Big AI knows it needs the content. Some of the companies concerned don't give a damn; money now, worry later is their motto. Those who can look past the revenue numbers are well aware they're engaged in what Microsoft itself called a “doom loop” of AI content strategy in those unsealed court documents. Will that stop them? Nah. According to the filing [PDF], which refers to sworn deposition testimony of OpenAI’s own corporate representative, it also testified that its LLMs were happy to vacuum up other people's work even when a paywall nominally protected it. Its rep was quoted as saying they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When OpenAI cofounder Greg Brockman was told OpenAI could hack its way through the firewall, he responded, “ah, nice.” Whether or not these statements are found to be reflective of a wider attitude, similar attitudes prevail amongst Big AI. Such policies could eventually come back and bite the LLM makers; it's already affecting the market for journalism, fiction writing, graphics creation, anything that demands humans actually get paid for producing original work. But Big AI just doesn't want to pay for it. That's not a side effect. It's the business model. Joe User couldn't care less. They just want a quick answer that sounds right. Some of them couldn't care less about getting the right answers. As for looking deeper to see if the information they're regurgitating is accurate, forget about it! As an OpenAI software engineer put it: “No matter how prominently we show the links, users won’t click.” Well, I click. That's a big reason why Perplexity is my AI of choice. It's not that it gives a more reliable answer. No, it's that, unlike most LLMs, Perplexity provides sources and links, and I check them before accepting what it tells me. But then I'm a journalist whose degrees are in history, where my professors drummed into me that you always - always - look at the primary sources. Big AI takes the same approach with its code generators. The US Ninth Circuit recently handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub lawsuit. The court ruled that the plaintiffs hadn't established a claim under one particular provision of the Digital Millennium

无论这些言论是否被认定为反映了更广泛的态度,类似的态度在大型 AI 公司中普遍存在。此类政策最终可能反噬大语言模型制造商;它已经影响到了新闻业、小说创作、图形设计等任何需要人类因创作原创作品而获得报酬的市场。但大型 AI 就是不想付费。这不是副作用,这是商业模式。

Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that. We also know that AI companies hate paying for any of it, obeying the licenses attached to it, or, God forbid, sharing their revenue with the companies and people who created the work in the first place. Recently, Big AI's default way of doing business: "Take it now, argue about legality later," has become more in your face than ever. Look, for example, at the copyright fight between The New York Times and OpenAI and Microsoft. According to 404 Media’s reporting on recently unsealed court documents - arguments from the plaintiffs that the court has not yet ruled on - Microsoft allegedly knew what it was doing when it was importing the internet willy-nilly. Microsoft's Director of Applied Science, Dr. Brent Hecht, was quoted in the news plaintiffs' 92-page combined brief as saying: the case was about “an astonishing theft of unprecedented proportions." Indeed, it was possibly the “largest theft of labor in human history,” he was quoted as saying in the summary judgment brief from the journalists' lawyers [PDF].This, mind you, wasn't a comment by a member of the press; it was from a senior Microsoft staffer. Hecht wasn't the only one at Microsoft who commented. In an internal Microsoft policy document also quoted in the court papers, the authors admitted generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain.” That, in turn, leads to model collapse. Ironically, AI is killing the content-creator goose that lays the gold eggs of worthwhile content. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not yet issued a decision on the case. OpenAI and Microsoft are actively disputing the plaintiffs' allegations under a "fair use" defense. They contend that using public articles and books to train large language models helps push forward public knowledge, and isn't acting as a "unlawful economic market substitute". But the statements the documents cites are real. Big AI knows it needs the content. Some of the companies concerned don't give a damn; money now, worry later is their motto. Those who can look past the revenue numbers are well aware they're engaged in what Microsoft itself called a “doom loop” of AI content strategy in those unsealed court documents. Will that stop them? Nah. According to the filing [PDF], which refers to sworn deposition testimony of OpenAI’s own corporate representative, it also testified that its LLMs were happy to vacuum up other people's work even when a paywall nominally protected it. Its rep was quoted as saying they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When OpenAI cofounder Greg Brockman was told OpenAI could hack its way through the firewall, he responded, “ah, nice.” Whether or not these statements are found to be reflective of a wider attitude, similar attitudes prevail amongst Big AI. Such policies could eventually come back and bite the LLM makers; it's already affecting the market for journalism, fiction writing, graphics creation, anything that demands humans actually get paid for producing original work. But Big AI just doesn't want to pay for it. That's not a side effect. It's the business model. Joe User couldn't care less. They just want a quick answer that sounds right. Some of them couldn't care less about getting the right answers. As for looking deeper to see if the information they're regurgitating is accurate, forget about it! As an OpenAI software engineer put it: “No matter how prominently we show the links, users won’t click.” Well, I click. That's a big reason why Perplexity is my AI of choice. It's not that it gives a more reliable answer. No, it's that, unlike most LLMs, Perplexity provides sources and links, and I check them before accepting what it tells me. But then I'm a journalist whose degrees are in history, where my professors drummed into me that you always - always - look at the primary sources. Big AI takes the same approach with its code generators. The US Ninth Circuit recently handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub lawsuit. The court ruled that the plaintiffs hadn't established a claim under one particular provision of the Digital Millennium

普通用户根本不在乎。他们只想要一个听起来正确的快速答案。有些人甚至不在乎答案是否正确。至于深入核实他们复述的信息是否准确,别想了!

Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that. We also know that AI companies hate paying for any of it, obeying the licenses attached to it, or, God forbid, sharing their revenue with the companies and people who created the work in the first place. Recently, Big AI's default way of doing business: "Take it now, argue about legality later," has become more in your face than ever. Look, for example, at the copyright fight between The New York Times and OpenAI and Microsoft. According to 404 Media’s reporting on recently unsealed court documents - arguments from the plaintiffs that the court has not yet ruled on - Microsoft allegedly knew what it was doing when it was importing the internet willy-nilly. Microsoft's Director of Applied Science, Dr. Brent Hecht, was quoted in the news plaintiffs' 92-page combined brief as saying: the case was about “an astonishing theft of unprecedented proportions." Indeed, it was possibly the “largest theft of labor in human history,” he was quoted as saying in the summary judgment brief from the journalists' lawyers [PDF].This, mind you, wasn't a comment by a member of the press; it was from a senior Microsoft staffer. Hecht wasn't the only one at Microsoft who commented. In an internal Microsoft policy document also quoted in the court papers, the authors admitted generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain.” That, in turn, leads to model collapse. Ironically, AI is killing the content-creator goose that lays the gold eggs of worthwhile content. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not yet issued a decision on the case. OpenAI and Microsoft are actively disputing the plaintiffs' allegations under a "fair use" defense. They contend that using public articles and books to train large language models helps push forward public knowledge, and isn't acting as a "unlawful economic market substitute". But the statements the documents cites are real. Big AI knows it needs the content. Some of the companies concerned don't give a damn; money now, worry later is their motto. Those who can look past the revenue numbers are well aware they're engaged in what Microsoft itself called a “doom loop” of AI content strategy in those unsealed court documents. Will that stop them? Nah. According to the filing [PDF], which refers to sworn deposition testimony of OpenAI’s own corporate representative, it also testified that its LLMs were happy to vacuum up other people's work even when a paywall nominally protected it. Its rep was quoted as saying they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When OpenAI cofounder Greg Brockman was told OpenAI could hack its way through the firewall, he responded, “ah, nice.” Whether or not these statements are found to be reflective of a wider attitude, similar attitudes prevail amongst Big AI. Such policies could eventually come back and bite the LLM makers; it's already affecting the market for journalism, fiction writing, graphics creation, anything that demands humans actually get paid for producing original work. But Big AI just doesn't want to pay for it. That's not a side effect. It's the business model. Joe User couldn't care less. They just want a quick answer that sounds right. Some of them couldn't care less about getting the right answers. As for looking deeper to see if the information they're regurgitating is accurate, forget about it! As an OpenAI software engineer put it: “No matter how prominently we show the links, users won’t click.” Well, I click. That's a big reason why Perplexity is my AI of choice. It's not that it gives a more reliable answer. No, it's that, unlike most LLMs, Perplexity provides sources and links, and I check them before accepting what it tells me. But then I'm a journalist whose degrees are in history, where my professors drummed into me that you always - always - look at the primary sources. Big AI takes the same approach with its code generators. The US Ninth Circuit recently handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub lawsuit. The court ruled that the plaintiffs hadn't established a claim under one particular provision of the Digital Millennium

一位 OpenAI 软件工程师曾说:“无论我们多么显眼地展示链接,用户都不会点击。” 好吧,我会点击。这正是 Perplexity 成为我首选 AI 的一个重要原因。并非因为它给出的答案更可靠。不,是因为与大多数大语言模型不同,Perplexity 提供来源和链接,我在接受它的回答前会去核查。但我毕竟是一名记者,我的学位是历史学,教授们反复灌输我们:永远——永远——要查阅原始资料。大型 AI 公司在代码生成器上采取了同样的做法。美国第九巡回法院最近在 Doe 诉 GitHub 案中判 GitHub、微软和 OpenAI 取得了一场有限的胜利。法院裁定,原告未能根据《数字千禧年》某一特定条款确立诉由

Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that. We also know that AI companies hate paying for any of it, obeying the licenses attached to it, or, God forbid, sharing their revenue with the companies and people who created the work in the first place. Recently, Big AI's default way of doing business: "Take it now, argue about legality later," has become more in your face than ever. Look, for example, at the copyright fight between The New York Times and OpenAI and Microsoft. According to 404 Media’s reporting on recently unsealed court documents - arguments from the plaintiffs that the court has not yet ruled on - Microsoft allegedly knew what it was doing when it was importing the internet willy-nilly. Microsoft's Director of Applied Science, Dr. Brent Hecht, was quoted in the news plaintiffs' 92-page combined brief as saying: the case was about “an astonishing theft of unprecedented proportions." Indeed, it was possibly the “largest theft of labor in human history,” he was quoted as saying in the summary judgment brief from the journalists' lawyers [PDF].This, mind you, wasn't a comment by a member of the press; it was from a senior Microsoft staffer. Hecht wasn't the only one at Microsoft who commented. In an internal Microsoft policy document also quoted in the court papers, the authors admitted generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain.” That, in turn, leads to model collapse. Ironically, AI is killing the content-creator goose that lays the gold eggs of worthwhile content. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not yet issued a decision on the case. OpenAI and Microsoft are actively disputing the plaintiffs' allegations under a "fair use" defense. They contend that using public articles and books to train large language models helps push forward public knowledge, and isn't acting as a "unlawful economic market substitute". But the statements the documents cites are real. Big AI knows it needs the content. Some of the companies concerned don't give a damn; money now, worry later is their motto. Those who can look past the revenue numbers are well aware they're engaged in what Microsoft itself called a “doom loop” of AI content strategy in those unsealed court documents. Will that stop them? Nah. According to the filing [PDF], which refers to sworn deposition testimony of OpenAI’s own corporate representative, it also testified that its LLMs were happy to vacuum up other people's work even when a paywall nominally protected it. Its rep was quoted as saying they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When OpenAI cofounder Greg Brockman was told OpenAI could hack its way through the firewall, he responded, “ah, nice.” Whether or not these statements are found to be reflective of a wider attitude, similar attitudes prevail amongst Big AI. Such policies could eventually come back and bite the LLM makers; it's already affecting the market for journalism, fiction writing, graphics creation, anything that demands humans actually get paid for producing original work. But Big AI just doesn't want to pay for it. That's not a side effect. It's the business model. Joe User couldn't care less. They just want a quick answer that sounds right. Some of them couldn't care less about getting the right answers. As for looking deeper to see if the information they're regurgitating is accurate, forget about it! As an OpenAI software engineer put it: “No matter how prominently we show the links, users won’t click.” Well, I click. That's a big reason why Perplexity is my AI of choice. It's not that it gives a more reliable answer. No, it's that, unlike most LLMs, Perplexity provides sources and links, and I check them before accepting what it tells me. But then I'm a journalist whose degrees are in history, where my professors drummed into me that you always - always - look at the primary sources. Big AI takes the same approach with its code generators. The US Ninth Circuit recently handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub lawsuit. The court ruled that the plaintiffs hadn't established a claim under one particular provision of the Digital Millennium