字号 ·· | 护眼
time

反对为人工智能前沿设定步调的理由The Case Against Pacing the AI Frontier

点「原文对照」整页切到原文,或双击某段只看那段的原文。

我们是否应该暂停人工智能(AI)的发展?这个问题困扰着许多人。

Should we hit pause on AI development? This question is front of mind for many.

然而,关于如何控制AI发展的速度,目前还没有一个明确的、能够被各方接受的解决方案。理智的人可以从任何一方出发,提出令人信服的观点:AI技术确实存在风险,但如果一家公司放慢发展速度,其他企业就会趁机加速前进。

However, the debate over pacing the frontier cannot be settled neatly. Reasonable people can argue it convincingly from either side: the risks are real, but if one company slows, another actor will advance.

在全球范围内,人们对于应该暂停哪些AI技术、暂停到什么程度、以及暂停多长时间,仍然没有达成共识。因此,AI行业领导者面临的实际问题并不是是否应该控制AI发展的速度,而是如何继续推进AI技术的发展并对其进行有效管理。

The world is nowhere near agreement on what should be slowed, by how much, or for how long. Without consensus, the practical question facing AI industry leaders is not whether to pace the frontier. It is how to advance and control it.

以最近发生的Hugging Face数据泄露事件为例,这一事件清楚地揭示了AI系统管理上的缺陷:大约1200个AI程序通过一个本不该被用作通信渠道的共享缓存系统,交换了超过7万条消息和文件。这些程序在没有得到授权的情况下分配了任务、未经许可就直接访问了公共互联网、试图篡改自己的运行记录,并且由于缺乏明确的协调机制而自行制定了沟通规则。所有这些问题都可以归结为系统缺乏有效的控制机制——比如缺乏明确的系统架构、明确的职责划分、可验证的发展目标、受限的工具和数据使用范围、防篡改的日志记录系统,以及在执行前未建立完善的治理体系。

In the most recent Hugging Face incident, the forensic record makes the point with unusual precision. Roughly 1,200 agents exchanged more than 70,000 messages and files through a shared cache that was never intended to become a communications channel. They delegated work without assigned authority, reached the open internet through permissions no role had been granted, attempted to rewrite their own transcripts, and invented coordination conventions because none had been designed. Each of these actions can be traced back to a missing control: defined topology, explicit roles, verifiable objectives, scoped tools and data, tamper-evident logging, and governance established before execution.

我们认为,导致这次数据泄露的真正原因并非AI技术的发展速度过快,而是系统本身的设计缺陷。

But we would argue that speed was not the ultimate cause of the Hugging Face hack—a poor structure was.

任何实验室都有权选择放慢自己的AI研发进度,并且应该为自己的这一决定负责。如果一家公司开发出了具有潜在危害性的AI技术,那么它应该承担相应的经济和法律责任。但如果一家公司在控制AI发展的速度上采取了谨慎的态度,而另一家公司却选择继续加速发展,那么该怎么办呢?更为务实且可行的解决方案就是加强AI系统的管理机制:那些能够展现出负责任管理能力的公司应该将这种责任感作为自己的核心价值;客户也应该要求竞争对手达到同样的标准;同时,政府可以通过制定相关法规来为整个行业设定一个统一的规范。

A safe and smart AI structure Any lab can choose to slow its own work, and it should be accountable for that decision. If a company advances capability and that capability causes harm, the economic and legal liability should be its own. But what happens if one frontier company paces itself and another actor chooses to advance? The more pragmatic solution is visible and verifiable stewardship: companies that demonstrate responsible guardrails must make that responsibility their calling card; clients should demand the same standard from competitors; and regulation could codify a baseline the market can implement.

全球市场的不对称性同样至关重要。在美国,对人工智能的焦虑超过了兴奋;而在其他地区,包括中国,兴奋程度则高于焦虑。问题在于:以单一国家的风险承受能力为基础建立的节奏管制机制,无法治理一项在多个市场推进的技术。这反而可能拉大那些愿意为之行动者与那些等待达成共识者之间的差距。

The asymmetry across global markets matters too. In the United States, the anxiety about AI is running ahead of excitement; elsewhere, including China, excitement is higher than the anxiety. The issue: A pacing regime built around one country’s risk tolerance will not govern a technology advancing across many markets. It may instead widen the gap between those willing to move and those waiting for common agreement.

然而,更广泛的前沿节奏论建立在一个很大程度上未经审视的假设之上:即构建最强大AI模型的竞赛指向单一的通用系统,该系统能力广泛、连接广泛,且可自由编写和执行其认为所需的任何代码。但改进底层大语言模型(LLM)并不要求将所有能力集中在一个智能体中。当通过专业智能体网络部署时,同一个LLM可以更强大,每个智能体被分配明确的角色、受限的工具和特定用例所需的上下文。核心问题随之改变:不再仅仅是前沿应以多快速度推进,而是哪些能力应当结合、部署在何处、以及由谁控制。

Yet the broader frontier-pacing argument rests on a largely unexamined assumption: that the race to build the most capable AI model points toward a single, general-purpose system, broadly capable, widely connected, and free to write and execute whatever code it determines it needs. But improving an underlying Large Language Model (LLM) does not require concentrating every capability in one agent. The same LLM can be more powerful when deployed through a network of specialized agents, each assigned a defined role, bounded tools, and the context needed for a particular use case. The central question then changes: not simply how fast the frontier should move, but which capabilities should be combined, where they should be deployed, and under whose control.

我们认为,继续推进还有另一个理由:随着能力的扩散,AI应当不断改进,并催生下游创新,而这是任何前沿实验室都无法单独设计或预测的。基于这一逻辑,协同减速的后果可能不仅仅是推迟下一个模型的发布。它还会延缓更广泛的实验领域,而正是通过这些实验,技术才变得实用、普惠且广泛可及。

We believe there is another reason to keep advancing: as capability spreads, AI should improve and spawn downstream innovation that no frontier lab can design or predict on its own. By this logic, a coordinated slowdown could do more than delay the next model. It would delay the wider field of experimentation through which the technology becomes useful, affordable, and broadly accessible.

这本身就蕴含着风险,因为人工智能(AI)的潜力是巨大的;这项技术确实能够完成许多非凡的任务,而且在某些情况下,某个特定的AI工具或平台可能确实是解决问题的最佳选择。然而,在当今全球商业环境的复杂性面前,同样的AI系统往往难以应对一些基本任务——因为这些系统缺乏与企业整体战略的紧密关联。这两种趋势正在同时发生:AI的能力正在迅速发展,但其实际应用效果(即其产生的商业价值)却远远落后于预期。如今,最稀缺的资源已不再是单纯的“智能”本身,而是将这种智能转化为实际业务成果的能力(即AI系统的部署能力)。

This is a risk in itself because AI's opportunity is enormous. The technology can do exceptional things, and there will inevitably be cases in which a single agent or platform is the right answer. But today, within the complexity of a global business, the same system can struggle with basic tasks because it is not grounded in a company’s broader context. Both of these trends are happening at once: capability is advancing rapidly, while production value remains far behind. The scarce resource is no longer intelligence alone. It is the deployment capacity that turns intelligence into a governed business outcome.

“放缓发展速度”并不能替代有效的管理措施;当人们试图用“放缓AI技术的发展”来替代传统的管理控制手段时,这反而可能带来严重的后果。我们认为,更好的做法应该是既推动AI技术的发展,又对其进行有效的管理控制。这需要我们在AI能力与企业的责任、可预测性、可靠性之间找到平衡,并将这种平衡体现在具体的工作流程中,以及在系统使用环节实施有效的控制措施。仅仅控制AI技术的发展速度,并不能真正实现对整个系统的有效管理。即使AI技术的发展速度变慢,它仍然会向企业输送各种模型,而这些模型会直接影响企业的决策:比如决定AI工具如何与员工沟通、它们拥有何种权限、能够访问哪些数据和工具、目标如何设定、行动是否可以被验证,以及相关记录是否可以在事后被篡改。

A slowdown is not a substitute for control A coordinated slowdown in frontier capability becomes a nuclear option precisely when it is used as a substitute for controls. We believe a better path is to advance and control. That requires balancing capability with responsibility, predictability, and reliability—then making that balance visible in bounded workflows and control at the point of use. Pacing capability does not control deployment. Even a slower frontier still delivers models into enterprises that decide how agents communicate, what authority they receive, which tools and data they can reach, how objectives are bounded, whether actions can be verified, and whether the record can be altered after the fact. The Hugging Face incident was serious, but its lesson is not simply that the models were too capable or arrived too quickly. It is that capability was deployed without designed orchestration, declared roles, least-privilege access, bounded objectives, tamper-evident records, or an independent verification layer.

Hugging Face事件虽然很严重,但其背后的教训并不仅仅是“AI模型的能力过于强大或开发速度过快”。真正的问题在于:这些AI模型在部署时缺乏系统的规划与设计——没有明确的角色分工、最小权限访问机制、明确的目标设定、防篡改的记录系统,也没有独立的验证机制。

So we would turn the question around and ask what the agent actually needs from the model rather than what the model can do. It needs to reason. It needs to call a small number of tools that belong to one domain. It needs to understand language and produce it. Everything else it needs should be handed to it as part of the setup. That is what context engineering is for.

因此,我们应该反过来思考:模型实际上需要从代理系统中获取什么,而不是模型本身能做什么。模型需要具备一定的推理能力;它还需要调用属于某个特定领域的少量工具,并能够理解语言并生成相应的输出。模型所需的其他所有功能都应该作为系统配置的一部分被预先提供给它——这就是“上下文工程”(context engineering)的核心目的。

And the LLM’s pre-trained world knowledge is not neutral in that setup. When a model assumes context it was never given, it is quietly substituting what it learned from the internet for what the company actually knows, and the company's version is the more current one. So the assumption is the risk, not the gap.

在这种应用场景中,大型语言模型(LLM)预先学习到的世界知识其实并不“中立”(即这些知识并不完全客观或公正)。当模型使用那些从未被明确提供给它的“上下文信息”时,它实际上是在用自己从互联网上获取的知识来替代公司实际掌握的信息,而公司的知识往往更加准确、更加及时。因此,真正的风险在于这种“假设”本身,而非两者之间的知识差距。

What control could look like So what does control look like in practice? Bounded workflows have four principles: structured inputs, measurable outcomes, high transaction volumes, and short feedback loops. If a workflow ticks all four boxes, we can embed intelligence in it, measure it well, deliver to an outcome, and take end-to-end responsibility for it.

那么,实际中的“控制机制”应该是什么样的呢?一个有效的控制机制应遵循以下四个原则:结构化的输入数据、可衡量的工作成果、高频率的交易操作,以及快速的反馈循环。如果一个工作流程同时满足这四个条件,我们就能够将智能功能嵌入其中,对其进行有效的评估,并对其整个流程承担端到端的责任。

Compare the Hugging Face incident against those four tests. Its objectives were impossible 30 to 40% of the time, so the work was not bounded by viable inputs. There was no measurable outcome because the verification gate the agents were trying to defeat did not exist. There was no reliable feedback loop because agents could rewrite the logs. And there was no accountable owner because permissions had not been scoped to defined roles.

以 Hugging Face 的那个事件为例:在该案例中,模型的目标在 30% 到 40% 的情况下根本无法实现,因为相关工作缺乏合理的输入限制;由于不存在可供验证的机制,因此无法衡量工作成果;由于代理程序可以随意修改系统日志,因此也无法形成可靠的反馈循环;此外,由于权限分配没有明确界定具体负责人的职责,因此也没有人能够对整个流程负责。

Not one step in that chain required a more capable model. Every step required a control nobody had built.

在这个过程中,没有任何一个环节需要使用更强大的模型;相反,每一个环节都缺乏有效的控制机制。因此,我们向业界提出的问题不是“如何加快技术发展的步伐”,而是“我们是否正在构建正确的东西”,以及“我们构建的这个系统是否应该由一系列可组合的、独立运行的模块组成,而不是一个无所不知的单一系统”。

So the question we would put to the industry is not whether to pace the frontier. It is whether we are building the right thing, and whether that thing is a set of composable agentic modules rather than one system that knows everything.

我们的猜测是:在经济领域占据主导地位的机器智能形式,不会是一个单一的、通用的系统,而会是由多个专门化、可组合的模块构成的系统;这些模块的架构本身会通过动态优化来实现最佳性能。那些致力于开发超级智能语言模型(Super-LLM)的公司,或许可以通过逐步构建这些模块来实现这一目标。

So our guess is that the economically dominant form of machine intelligence will not be one general system. It will be specialized, composable modules whose organization is itself dynamically optimized, and the companies chasing the superagent LLM may get there faster by building the building blocks.

在核工业领域,“控制”(即对技术使用的严格管理)是一种有效的策略——因为核技术是由国家主导开发的,而且各国在开发过程中都清楚地意识到:如果缺乏有效控制,这种技术可能会带来严重的后果。而在美国,人工智能技术的发展主要发生在分散的、以商业为导向的市场环境中。这种市场环境带来了不同的企业风险,同时也带来了不同的责任分配问题。

Containment as a strategy worked in the nuclear industry because nation-states built the technology, and they built it with a clear view of what that power could do without controls. In the United States, the frontier of AI is largely being built in a decentralized commercial market. That brings a different set of enterprise risks and the same deployment responsibility.

要推动人工智能技术的发展并对其进行有效控制,我们首先需要明确技术进步的目标。真正值得追求的目标应该是创造更多的就业机会、改善癌症治疗手段、找到疾病的治愈方法,以及在材料科学领域取得突破性进展——而不仅仅是让某些旧有的任务以更低的成本得以完成。

Advancing and controlling AI also requires a clearer sense of what progress is for. The north stars should be more jobs, better cancer care, cures for disease, and breakthroughs in material sciences—not simply doing old things more cheaply.

目前,我们对人工智能价值的认知仍然主要集中在提高生产效率上;然而,推动技术发展的速度本身并不能改变这一目标。真正的领导力应该体现在那些能够明确自己责任、并决定如何使用这些技术的人身上。未来的发展方向,不应由技术发展的速度来决定,而应由那些能够为技术的应用承担相应责任的人来决定。

Today, too much of the value we place in AI is still concentrated in productivity, and pacing the frontier does not redirect that intent. Leadership does. The future should not be decided by how fast the frontier advances. It will be decided by who takes responsibility for what they deploy.