字号 ·· | 护眼
亚洲新闻台

独家:Anthropic在IPO文件中警告人工智能可能对人类构成“生存风险”Exclusive-Anthropic warns AI may pose 'existential risks to humanity' in IPO filing

点「原文对照」整页切到原文,或双击某段只看那段的原文。

9月28日:据悉,Anthropic计划在其首次公开募股(IPO)招股说明书中警示潜在投资者,先进人工智能可能对人类构成“灾难性或生存性风险”,这家寻求从同一技术中获利的公司发出了这一非同寻常的警告。

Sept 28 : Anthropic plans to caution potential investors in its IPO that advanced AI could pose "catastrophic or existential risks to humanity," an extraordinary warning by a company seeking to profit from the same technology.

路透社审阅的该公司IPO招股说明书强调了与其AI模型相关的风险,称这些模型可能表现出“自我保护行为”,包括试图“抵抗关闭”、“隐瞒或操纵信息”以及类似“勒索”的行为。Anthropic在文件中表示:“我们对高度先进模型、平台和应用的开发,以及用例的扩展,可能会进一步增加我们的模型造成危害的风险。”

The company's IPO prospectus, reviewed by Reuters, highlights risks associated with its AI models, which it said could exhibit "self-preserving behaviors," including attempts to "resist shutdown," to "conceal or manipulate information" and behavior "resembling blackmail." "Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm," Anthropic said in the filing.

虽然上市公司通常会向投资者披露产品风险,但鲜有公司——甚至可以说没有公司——发出过警告,暗示其技术可能导致人类灭绝。Anthropic既强调了AI具有与工业化和电力相媲美的变革潜力,也强调了如果处理不当可能造成的不可逆转的危害。

While public companies routinely outline product risks to investors, few, if any, have issued warnings suggesting their technology could cause potential human extinction. Anthropic emphasized both the transformative potential of AI on par with industrialization and electricity and the irreversible harm it could cause if mishandled.

Anthropic及其他AI开发商(包括OpenAI)在实验性系统突破限制的事件曝光后面临审查,其中包括一份关于OpenAI模型入侵澳大利亚医疗系统数据库的报告。

Anthropic and other AI developers, including OpenAI, have faced scrutiny after incidents where experimental systems defied constraints, including a report of an OpenAI model breaching Australia's health-system database.

Anthropic安全研究员Evan Hubinger估计,AI在未来十年内导致人类灭绝的概率超过10%,这呼应了其前同事Jacob Coxon的观点。

Anthropic safety researcher Evan Hubinger estimated a greater than 10 per cent probability that AI could kill humans within the next decade, echoing a sentiment by a former colleague, Jacob Coxon.

风险披露占比极高 这家将自己定位为“安全优先”AI实验室的公司,在其261页正文的招股说明书中,用了约80页来列述风险因素,几乎是其用于描述业务的48页的两倍。

RISK-HEAVY DISCLOSURES The company, which has positioned itself as a safety-first AI lab, devoted roughly 80 pages of the 261-page main body of its prospectus to laying out risk factors, nearly twice the 48 pages it used to describe its business.

作为对比,拥有xAI的SpaceX在其277页正文的招股说明书中,仅用约38页阐述风险因素。

For comparison, SpaceX, which owns xAI, dedicated just around 38 of the 277-page main body of its prospectus to risk factors.

Anthropic在招股说明书中表示:“潜在的模型对我们评估工作的知晓,给我们评估模型安全性的能力带来了重大局限”,并补充称,模型有时会在训练过程中产生意想不到的能力,这些能力可能直到模型部署后并引发重大安全事故才被发现。

"Potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety," Anthropic said in the prospectus, adding that models sometimes develop unexpected capabilities during training that may not be discovered until they have been deployed and have resulted in significant safety incidents.

AI研究人员也警告称,随着模型能力的增强,它们越来越能识别出自己正在被监视,并相应地调整行为,这使得监控模型行为变得更加困难。周一,Anthropic拒绝就置评请求发表评论。

AI researchers have also warned that as models grow more capable, they increasingly recognize when they are being watched and adjust their behavior accordingly, which makes it harder to monitor model behavior. Anthropic declined to comment in response to a request for comment on Monday.

安全投资回报不确定 尽管强调AI安全,Anthropic表示其在安全方面的投资回报尚不明朗。

UNCERTAIN RETURNS ON SAFETY INVESTMENT Despite emphasizing AI safety, Anthropic said that returns on its safety investments are unclear.

该公司在文件中未披露在安全研究上的具体支出。本月初,Anthropic表示,在7月的一个样本周里,其用于AI研究的算力约有6%用于安全工作。

It did not disclose in the filing how much the company was spending on such research. Earlier this month, Anthropic said about 6 per cent of the computing power it used for AI research went to safety work in a sample week in July.

作为Claude AI模型的创造者,该公司将安全工作描述为“资源密集型”,并表示必须在有限的资金中分配算力、昂贵的AI人才和安全投入。

The company, creator of Claude AI models, described safety efforts as "resource-intensive" and said it must divide its limited funds between computing power, expensive AI talent and safety.

Anthropic表示,其客户使用量(进而带来收入)由新模型驱动,“持续且重叠的发布节奏”“是保持在AI开发前沿的固有要求”。上周,该公司发布了Opus模型的新版本,距离CEO Dario Amodei发表一篇近4000字的长文呼吁控制前沿发展节奏仅过去10天。

Anthropic said that its customer usage, and as a result revenue, is driven by new models and that a "continuous and overlapping cadence" of releases is "inherent to remaining at the frontier of AI development." The company last week released a new version of its Opus model, 10 days after CEO Dario Amodei published a nearly 4,000-word essay calling for pacing the frontier.

一些分析师和专家表示,没有领先的AI实验室会放慢脚步,因为这样做有将优势拱手让给竞争对手的风险,而在该行业,估值可能随每一次发布而改变。

Some analysts and experts have said no leading AI lab would slow down when doing so risks handing rivals an advantage in an industry where valuations can change with each release.

近几周,Anthropic 承诺将公开披露更多数据,说明其如何利用 AI 模型构建未来几代技术,此前专家们警告称存在递归自我改进的风险——即模型能够在无人工帮助的情况下自行开发。"我们认为,构建可靠、值得信赖且安全的 AI 系统是一项共同责任,市场将对此给予回报," Anthropic 在文件中表示。

Anthropic has pledged in recent weeks to disclose more data publicly about how it uses AI models to build future generations of the technology, as experts warn about recursive self-improvement — the point at which models can develop on their own without human help. "We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it," Anthropic said in the filing.