字号 ·· | 护眼
连线

这些人工智能专家希望公开进行高风险研究These AI Experts Want to Do High-Stakes Research Out in the Open

点「原文对照」整页切到原文,或双击某段只看那段的原文。

行业科学家Nathan Lambert和Tom Zick持相反观点。二人创立了非营利组织Trillium Labs,将以更透明的方式开展各领域AI研究——包括递归自我改进(RSI)和智能体等潜在问题领域。具体而言,这意味着公开实验细节,以便外部科学家能够研究和复现。

Nathan Lambert and Tom Zick, two industry scientists, believe the opposite. The pair founded a nonprofit, Trillium Labs, that will work on various areas of AI research—including potentially problematic areas like recursive self-improvement (RSI) and agents—in a more transparent way. In practice, this will mean publishing the details of experiments so that outside scientists can study and replicate them.

Lambert表示,前沿AI实验室对工作的保密降低了社区审查想法和贡献新方法的能力。他认为,让外部专家了解模型的构建和调优过程,可能对降低风险至关重要。

Lambert says the way frontier AI labs keep their work secret reduces the community’s ability to scrutinize ideas and contribute new approaches. He believes that letting outside experts see how models are built and tuned could be crucial to mitigating risks.

“过去几千年里,人类将科学方法纳入工具箱,作为减轻危害、构建更美好未来的途径,”Lambert告诉《连线》(WIRED)。“当前前沿AI发展的封闭轨迹正让我们倒退一步。”世界上最强大的模型,如OpenAI和Anthropic的模型,只能通过应用程序或应用程序编程接口访问。这往往以牺牲关于模型如何构建及如何表现的透明度为代价。

“Over the past few millennia, humanity has had the scientific method in our toolbox as a way to mitigate harms and build better futures,” Lambert tells WIRED. “The current closed trajectory of frontier AI development is taking us a step backwards.”The world’s most powerful models, like those from OpenAI and Anthropic, can only be accessed through an app or an application programming interface. Often, this comes at the cost of transparency about how the model is built and how it behaves.

其他公司,尤其是中国公司,提供相对强大的可下载模型,用户可在自有硬件上运行。例如,中国公司小米最近公开了其某模型大规模训练运行的实时细节。斯坦福大学研究人员正在开放预训练AI模型Marin。

Other companies, especially those in China, offer relatively powerful models that can be downloaded and run on a user’s own hardware. The Chinese company Xiaomi, for example, recently published live details of a major training run involving one of its models. And researchers at Stanford are pretraining the AI model Marin in the open.

行业目前陷入关于哪种策略最优的争论,主要因为前沿模型现已极其强大。它们能自动化发现新软件漏洞,并自动探测和入侵系统,近期高调黑客攻击事件更引发了更严格的审视。

The industry is currently locked in a battle over which strategy is best, mostly because of how powerful frontier models now are. They can automate the discovery of new software vulnerabilities and automatically probe and hack into systems, and recent high-profile hacking sprees have prompted even greater scrutiny.

有限访问制度的支持者认为,将这种权力掌握在少数受信任者手中至关重要,而兰伯特和齐克所在的阵营则认为,对风险的共同认知会让我们所有人都受益。

Proponents of a limited-access system say it’s crucial to keep that power in the hands of a trusted few, while those in Lambert and Zick’s camp believe that a shared understanding of the risks means we’re all better off.

兰伯特曾在Ai2工作,这是一个采取异常开放方式进行AI研究的实验室,包括在发布模型的同时公开构建模型所用的数据和训练方法的细节。他曾在Hugging Face任职,运营一个广受欢迎的技术博客,并创立了“美国真正开放模型”倡议,旨在鼓励美国公司发布更多开放模型。齐克曾在哈佛大学工作,并协助查尔斯·施瓦布制定“负责任AI”相关政策。

Lambert previously worked at Ai2, a research lab that has taken an unusually open approach to AI, including publishing details of the data and the training methods used to build models alongside the models themselves. He previously worked at Hugging Face, runs a popular technical blog, and founded the American Truly Open Models, an initiative aimed at encouraging US companies to release more open models. Zick worked at Harvard University and helped Charles Schwab devise policies around “responsible AI.”

两人在新冠疫情期间通过Zoom相识,当时两人均为加州大学伯克利分校从事AI研究的研究生。他们创办这家新非营利组织的想法,源于看到产业界AI研究与学术界工作变得多么脱节;兰伯特表示,教授和学生往往无法复现大公司实验室内部的工作,因为他们缺乏所需的资源。

The two met over Zoom during the COVID-19 pandemic, when both were graduate students at UC Berkeley working on AI. They got the idea for the new nonprofit after seeing how disconnected industry AI research has become from academic work; Lambert says professors and students are often unable to replicate the work going on inside big company labs because they lack the resources required.

齐克表示,今日启动的Trillium Labs将最初专注于后训练——即在大模型构建完成后对其进行微调。另一个关键领域将是RSI,这是一种让AI参与研究以开发新模型的过程。持续进步可能无限延续、导致人类失去控制的前景,让许多AI研究人员感到担忧。本月初,一位Anthropic研究人员离职并警告RSI可能对人类构成生存威胁,该问题因此获得主流关注。

Zick says Trillium Labs, which launched today, will initially focus on post-training—fine-tuning large models after they’ve been built. Another key area will be RSI, a process for developing new models by having AI contribute research. The prospect that ongoing progress could continue indefinitely, leading to a loss of human control, has alarmed many AI researchers. The issue gained mainstream attention earlier this month when an Anthropic researcher left the company and warned that RSI could pose an existential threat to humankind.

这家非营利组织还将研究强化学习——即奖励模型的良好结果、惩罚其不良结果——如何提升模型能力。这种方法让智能体变得更加强大,但也更倾向于做出意想不到的事情。他们将研究强化学习如何塑造AI模型的“性格”与行为,这种方法可能带来问题,例如当模型变得过度阿谀奉承时。

The nonprofit will also look at how reinforcement learning, which rewards a model for good results and punishes it for bad outcomes, can improve its capabilities. That approach has made agents far more capable, but also more inclined to do unexpected things. They’ll study how reinforcement learning shapes the character and behavior of AI models, a method that can pose problems when a model becomes overly sycophantic, for example.

“要理解强化学习在后训练阶段如何规模化,你需要大量算力和大量细致的实验,”Zick表示。她指出,公开强化训练运行的细节可能会产出令人惊讶的见解,因为外部研究人员会对这些工作进行审视。

“To understand something like how reinforcement learning scales in post-training, you need significant compute and a lot of careful experimentation,” Zick says. She says that publishing details of how reinforcement training runs work could yield surprising insights as outside researchers scrutinize the work.

该实验室已从Schmidt Sciences、Halcyon Futures等机构筹集了一笔未披露金额的资金。创始人表示,他们的目标是总共筹集4000万至1亿美元,并计划在未来18个月内投入3000万美元用于训练。

The lab has raised an undisclosed sum from Schmidt Sciences, Halcyon Futures, and others. The founders say they aim to raise $40 to $100 million in total and plan to spend $30 million on training over the next 18 months.

进步研究所(Institute for Progress,一个政策智库)新兴技术政策主任Tim Fist告诉《连线》(WIRED):“我非常支持比目前研发领域更高的透明度。”

“I'm a massive fan of much more transparency than we currently have in R&D,” Tim Fist, director of emerging technology policy at the Institute for Progress, a policy thinktank, tells WIRED.

Lambert和Zick最终希望Trillim Labs能为关于如何最好地构建AI的广泛讨论贡献一些亟需的细微差别。

Lambert and Zick ultimately hope that Trillim Labs will contribute some much-needed nuance to the wider discussion about how best to build AI.

“我们正处于一个由少数世界观主导AI话语的时代,”Lambert说。“我们相信,科学方法和对近期事件的严谨测量,是理解AI模型新行为的最佳途径。”

“We’re in an era of AI discourse dominated by a few world views,” Lambert says. “We believe that the scientific method and careful measurement of recent events is the best way to understand new behaviors of AI models.”