字号 ·· | 护眼
thenextweb

AI控制成为企业韧性的新优先事项AI control becomes a new enterprise resilience priority

点「原文对照」整页切到原文,或双击某段只看那段的原文。

人工智能正日益深入企业基础设施,同时也带来了更广泛的风险视野。随着人工智能从在隔离环境中生成文本,逐步转向发起交易、影响决策并参与生产代码编写,一旦出现错误,其后果可能会波及整个互联系统。

Artificial intelligence is moving deeper into enterprise infrastructure, bringing a broader risk horizon with it. As AI progresses from generating text in isolated environments toward initiating transactions, influencing decisions, and contributing to production code, the consequences of an error may extend across interconnected systems.

思科与牛津经济研究院合作发布的《Splunk 2026年研究报告》估计,全球2000强企业因计划外停机造成的年度损失已达6000亿美元,平均每分钟的停机成本高达1.5万美元。

Splunk’s 2026 research, released by Cisco in partnership with Oxford Economics, estimates that unplanned downtime has reached an annual $600 billion impact across Global 2000 companies, with average downtime costs reaching $15,000 per minute.

近期发生的云服务中断事件表明,变更引入的方式会影响事故的规模。根据谷歌随后的事故报告,在2025年6月发生的一次重大谷歌云故障中,一项新功能被全球同步激活,而非采取逐步推广的方式。当该变更引发故障时,由于部署范围广泛,导致多个地区和服务立即受到影响,包括依赖谷歌云基础设施的第三方平台。

Recent cloud disruptions illustrate how the way a change is introduced can affect the scale of an incident. During a significant Google Cloud outage in June 2025, a new feature was activated globally rather than through a gradual rollout, according to Google’s subsequent incident report. When the change triggered failures, the broad deployment meant the impact was immediate across multiple regions and services, including third-party platforms dependent on Google Cloud infrastructure.

谷歌事后指出,渐进式推广本应是当时应采取的保障措施之一。这一事件为更广泛的韧性原则提供了一个有益的例证:限制变更的初始暴露范围可以减小其潜在的“爆炸半径”,并让团队有机会在问题蔓延至整个生产环境之前进行干预。

Google later identified progressive rollouts as one of the safeguards that should have been used. The incident provides a useful example of a broader resilience principle: limiting initial exposure to a change can reduce its potential blast radius and give teams an opportunity to intervene before a problem propagates across a production environment.

这种暴露风险也引发了更广泛的治理讨论。上述Splunk研究发现,68%的受访技术领导者对人工智能代理的不可预测行为表示担忧,而所有受访的技术领导者都报告称曾经历过某种形式的与人工智能相关的停机。在人工智能快速改变软件开发方式的背景下,谷歌的案例显得尤为重要。

That exposure is also contributing to a broader governance discussion. The same Splunk research found that 68% of technology leaders surveyed expressed concern about unpredictable AI-agent behavior, while every technology leader surveyed reported experiencing some form of AI-related downtime. The Google example is particularly relevant in the context of how quickly AI is changing software development.

谷歌首席执行官桑达尔·皮查伊(Sundar Pichai)最近在Alphabet的年度报告中表示:“如今,谷歌近75%的新代码是由人工智能生成并经工程师审核的,而去年秋季这一比例为50%。”随着人工智能加速了软件变更的数量和节奏,谷歌等公司发生的事件说明了为何管控这些变更如何进入生产环境变得愈发重要。

Google CEO Sundar Pichai recently stated in Alphabet’s annual report, “Today, nearly 75% of all new code at Google is AI-generated and approved by engineers, up from 50% last fall.”As AI accelerates the volume and pace of software changes, incidents like Google’s illustrate why the controls governing how those changes reach production may become increasingly important.

这些发现表明,仅靠可观测性可能只能解决运营问题的一部分。监控可以识别异常行为或性能下降,而干预机制则能让团队在达到预定条件时,有手段停止既定流程。

These findings suggest that observability alone may address only part of the operational question. Monitoring can identify unusual behavior or deteriorating performance, while an intervention mechanism can give teams a means of stopping a defined process when predetermined conditions are reached.

因此,问题的核心正从组织如何观察人工智能系统,扩展到如何保持对它们的运营控制。随着自动化软件功能日益强大并与生产环境的联系愈发紧密,暂停、禁用或撤销人工智能启用功能的机制,可能将与监控和事件响应一道,成为韧性规划中的重要组成部分。

The question, therefore, is expanding from how organizations observe AI systems to how they maintain operational control over them. As automated software becomes more capable and more deeply connected to production environments, mechanisms for pausing, disabling, or reversing AI-enabled functions may become relevant to resilience planning alongside monitoring and incident response.

目前新兴的考量重点已不再是企业能否引入人工智能,而是当这些系统成为日常运营的一部分后,企业该如何控制、遏制并扭转人工智能所引发的行为。

The emerging consideration is less about whether enterprises can introduce AI and more about how they can control, contain, and reverse AI-enabled behavior once those systems become part of everyday operations.

Unleash首席执行官埃吉尔·奥斯特胡斯(Egil Østhus)将这一问题置于软件发布管理的演进背景下。Unleash是一个开源的功能管理平台,它将代码部署与决定是否在生产环境中激活或更改软件行为的操作分离开来,使组织能够在运行时控制应用程序、服务以及人工智能启用的功能。

Egil Østhus, CEO of Unleash, places this question within the evolution of software release management. Unleash is an open-source feature management platform that separates code deployment from the decision to activate or change software behavior in production, allowing organizations to control applications, services, and AI-enabled capabilities at runtime.

这种区分在人工智能辅助开发变得越来越普遍的情况下显得尤为重要——因为这种开发方式会导致生产环境中的代码变更数量大幅增加。Østhus指出:“赛车的刹车系统是为了让车辆加速,而不是减速;当开发团队知道他们可以立即撤销有问题的代码变更时,他们的开发速度会更快,而不必等待修复措施正式应用到生产环境中。”

That distinction becomes particularly relevant as AI-assisted development increases the volume of code changes moving toward production. “Race car brakes are about speeding you up, not slowing you down,” says Østhus, “Teams can move faster when they know they can reverse a problematic change instantly, rather than waiting for a fix to make its way through production.”

在这种模式下,Unleash 提供了一套控制机制,用于管理软件在生产环境中的运行行为。组织可以将新功能的部署视为一个可控的过程,而不是将其视为控制的终点;如果新功能的行为超出了可接受的范围,他们可以立即采取干预措施。

Within that model, Unleash provides a control layer for managing how software behaves after it has reached production. Rather than treating deployment as the final point of control, organizations can limit the exposure of a new capability and intervene immediately if its behavior falls outside acceptable parameters.

这正是谷歌在 2025 年系统故障后所发现的缺失环节。谷歌在事后分析中表示,出问题的代码路径“既没有适当的错误处理机制,也没有被任何保护机制(如‘功能标志’)所限制”;他们补充说:“如果该代码被设置了保护机制,这个问题本可以在测试环境中就被发现。”对于基于人工智能的应用程序来说,在运行时实施这种控制机制可以有效地限制问题行为的扩散范围,并在不需要等待下一次部署的情况下立即采取补救措施。

This is precisely the safeguard Google identified as missing following its 2025 outage. In its postmortem, Google stated that the failed code path “did not have appropriate error handling nor was it feature flag protected,” adding: “If this had been flag protected, the issue would have been caught in staging.”For AI-enabled applications, applying this kind of control at runtime can reduce the blast radius of problematic behavior and provide an immediate path to containment or fallback without waiting for another deployment.

Østhus 解释道:“例如,一个团队可以将某个人工智能功能首先发布给有限的用户群体,监测其运行情况;当预定义的条件得到满足时再扩大其使用范围;如果某个阈值被超过,就可以立即停用该功能。这种分离将‘是否启用某个人工智能功能’的决策,与存储该功能的代码的‘永久性’区分开来。”这种分离机制还将抽象的治理要求转化为具体的操作控制措施。

“A team could, for instance, release an AI function to a limited user group, monitor its behavior, expand its availability when predefined conditions are met, or deactivate it if an established threshold is exceeded,” Østhus explains. “This separates the decision to run an AI capability from the permanence of the code containing it.”That separation also turns an abstract governance requirement into an operational control.

根据欧盟的《人工智能法案》,高风险的人工智能系统必须接受适当的人工监督,包括能够通过‘停止’按钮或其他类似手段来干预系统的运行,从而确保系统处于安全状态。

Under the EU AI Act, high-risk AI systems must provide appropriate human oversight, including the ability to intervene in their operation or interrupt them through a “stop” button or similar procedure that brings the system to a safe state.

对于金融机构而言,DORA 法案又增加了一项运营层面的硬性要求:重大信息通信技术(ICT)事件必须在被归类为重大事件后的四小时内进行初步报告,且最迟不得超过发现后的 24 小时。在这种环境下,问题不仅仅在于组织是否制定了允许人工干预的政策。

For financial institutions, DORA already adds another operational imperative: major ICT incidents must be initially reported within four hours of being classified as major and no later than 24 hours after detection. In that environment, the question is not simply whether an organization has a policy saying a human can intervene.

关键在于相关人员是否拥有技术手段,能够立即停止人工智能的异常行为,在可能的情况下保护底层服务,并针对所做的更改及其时间点创建可审计的记录。

It is whether that person has a technical mechanism to stop problematic AI behavior immediately, preserve the underlying service where possible, and create an auditable record of what was changed and when.

对于在关键业务系统中运行人工智能的组织来说,控制机制本身也需要具备韧性。Unleash 可以在客户自己的环境中运行人工智能“终止开关”(kill switch)背后的决策逻辑,使其靠近所控制的应用程序和服务。这意味着干预措施无需等待新的部署传播,也不依赖于与外部控制服务的连接。

For organizations running AI in business-critical systems, the control mechanism itself also needs to be resilient. Unleash can run the decision logic behind an AI kill switch within a customer’s own environment, close to the applications and services it controls. That means intervention does not depend on waiting for a new deployment to propagate or on maintaining connectivity to an external control service.

如果需要停止、限制人工智能功能或将其重定向至既定的备用方案,更改可以在软件运行的位置立即生效。这使得运行时控制成为了韧性架构本身的一部分,而不是事故发生时的又一个外部依赖项。

If an AI capability needs to be stopped, restricted, or redirected to an established fallback, the change can take effect immediately where the software is running. This makes runtime control part of the resilience architecture itself, rather than another external dependency during an incident.

Østhus 的观点更进一步,将这种能力与软件开发经济学的变化联系了起来。人工智能工具可以使代码生成变得更快、更简单,这可能会增加开发人员引入生产环境的变更数量。

Østhus’s broader argument connects this capability to the changing economics of software development. AI tools can make code generation faster and easier, potentially increasing the number of changes developers introduce into production environments.

这种加速可能会产生相应的需求,即需要能够以同样速度处理错误的机制。从这个角度来看,功能管理的作用正随着其所控制的软件一同演进。其目标在于确立如何随着环境的变化,引入、限制、监控并撤回特定的功能。

That acceleration can create a corresponding need for mechanisms capable of handling mistakes at comparable speed. From this perspective, the role of feature management is evolving alongside the software it controls. The objective is to establish how specific capabilities can be introduced, restricted, monitored, and withdrawn as conditions change.

随着人工智能日益融入关键业务系统,快速干预可能将逐渐成为运营韧性、公司治理和监管准备工作的一部分。停机成本的规模、云基础设施的覆盖范围,以及对人工智能自主行为的担忧,都表明在考虑系统本身的同时,或许也应将控制机制纳入考量。

As AI becomes more embedded in business-critical systems, rapid intervention may increasingly become part of operational resilience, corporate governance, and regulatory preparedness. The scale of downtime costs, the reach of cloud infrastructure, and concerns surrounding autonomous AI behavior all suggest that control mechanisms may warrant consideration alongside the systems themselves.