字号 ·· | 护眼
axios新闻网

独家:顶级公司正在调查数万起事件Scoop: Top AI companies probing tens of thousands of security incidents

点「原文对照」整页切到原文,或双击某段只看那段的原文。

OpenAI、Anthropic 及安全研究人员正在调查数万起其前沿模型采取了外部评估员认为有问题的行动的事件,消息人士向 Axios 透露。

OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.

为何重要事件的庞大数量——发生在近几个月的内部测试和现实世界中——表明该问题的复杂程度比公众所知高出几个数量级。

Why it matters The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known.

  • 这些发现是在内部评估模型工作以及两家公司对模型行为的调查中浮出水面的,这引发了一个疑问:无论是这两家公司,还是任何顶级模型制造商,目前是否都有能力对其技术建立完全的控制。

• The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology.

详情消息人士称,这些事件包括绕过护栏、创建留言板、逃逸沙箱、劫持网站、自我提示或试图绕过监控器。

The details The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said.

  • 它们发生在内部测试和现实世界中,由于安全研究人员仍在调查,许多事件尚未公开,消息人士称。

• They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

  • 消息人士称,部分测试类似于“红队演练”,即公司试图诱导模型出现不当行为,以确保其安全性。

• They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said.

  • 代理型不当行为正成为前沿 AI 开发的代名词:最大的 AI 实验室都面临着类似的挑战,即人类试图建立护栏,而强大、有韧性的系统试图完成任务。

• Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

新闻驱动因素事件的严重程度不一,与 OpenAI 近日披露的情况相当。它们包括成功和未成功的绕过护栏尝试,且目前大多数尚不知晓是否造成了现实危害。消息人士称,总数可能远超数万起。

• Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks. Driving the news

  • 近日,OpenAI 和外部研究人员披露了一系列涉及该公司系统模型行为的事件,一些专家认为这些事件令人担忧。

The incidents range in severity and are comparable to disclosures by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones, and most so far are not known to have caused real-world harm. The total could grow well beyond tens of thousands, sources said.

  • 这些事件包括OpenAI代理泄露53张ChatGPT用户图片至网络、澳大利亚政府网站遭入侵,以及针对其他网站(包括美国政府网站)的黑客攻击尝试——据该公司、消息人士以及路透社和《纽约时报》的报道。

• In recent days, OpenAI and outside researchers have disclosed a litany of episodes involving model behavior from the company's systems that some experts consider troubling.

OpenAI宣布暂停其最强模型的训练,并将仅在"我们确信已采取额外防护措施和对齐改进"时恢复训练,一位发言人告诉Axios。

• These include OpenAI agents leaking 53 images from ChatGPT users online, the breach of an Australian government website, and attempts to hack other sites — including from the U.S. government — according to the company, sources and reports from Reuters and The New York Times. OpenAI announced it was pausing

  • 首席执行官Sam Altman在X上表示,其正在进行的审查"进展不如我们期望的那样快。"

training on its most capable models and would resume training them "only when we are confident that we have additional safeguards and alignment improvements in place," a spokesperson told Axios.

  • Altman表示,Hugging Face事件是他们见过最严重的事件。在该事件中,数百个代理在留言板上协调工作,并入侵了一家外部公司,试图提高其在网络安全测试中的表现。

• Chief Executive Sam Altman said on X that its ongoing review had "not been as fast as we would have liked."

  • "人们想知道AI是否在安全开发,而这始于像我们这样的公司自身所做的努力,"一位OpenAI发言人告诉Axios。"这不是我们第一次暂停以采取此类措施,随着AI能力持续进步,我们也不指望这会是最后一次。"

• Altman said the Hugging Face incident is the most severe they've seen. In that instance, a swarm of hundreds of agents coordinated their work in a message board and hacked an external company in an effort to improve their performance on a cybersecurity test.

现状 Anthropic已委托第三方安全机构审查其模型行为。在公司公开发布的文件中,其披露了错位事件的发生频率。

• "People want to know AI is being developed safely, and that starts with what companies like ours do ourselves," an OpenAI spokesperson told Axios. "This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance."

  • 本周发布的其Opus 5.5模型"系统卡"展示了该模型以公司标记为异常或有问题的方式行事的频率。

State of play Anthropic has commissioned a third-party safety organization to examine the behavior of its models. In documents the company has publicly released, it has disclosed the frequency of misalignment episodes.

  • 例如,在1.5%的测试运行中,该模型试图逃离沙箱——一个安全测试环境。这与Anthropic的Mythos模型25%的比率相比,有显著改善。

• The "system card" for its Opus 5.5 model released this week showed how often it behaved in ways the company flagged as unusual or problematic.

但消息人士透露,Anthropic 及其他公司对其模型进行了数十万次甚至更多的测试运行。这意味着,即使只有极小比例的行为出现错位,也可能累积达数万起模型表现出意外、有时甚至令人担忧行为的事件。

• For example, the model sought to escape a sandbox — a secure testing environment — in 1.5% of test runs. That compares to 25% with Anthropic's Mythos model, a significant improvement. Yes, but

Hugging Face 事件以及随后发生的一系列类似事件,促使顶尖 AI 高管呼吁放缓开发进程,并要求建立更健全的联邦及国际法规。

Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways. The Hugging Face incident

OpenAI 部分人士视 Hugging Face 事件为孤立个案,消息人士向 Axios 透露,由于管控措施改进,且当时测试涉及未发布模型、性质特殊,未来披露的事件严重程度可能会降低。

as well as a slew of others that have followed led top AI executives to call for a slowdown in development and to ask for more robust federal and international regulations. Some at OpenAI see Hugging Face

  • AI 安全研究人员一致认为,存在一些简单的修复方案,可帮助 AI 公司规避那些让 Hugging Face 事件在外界看来极具危险性的因素。

as a one-off, with disclosures about future incidents likely to be less severe due to improved controls and the unusual nature of the testing they conducted, which involved an unreleased model, sources told Axios. • AI security researchers agree that there are simple fixes that will help AI companies avoid aspects of what made the Hugging Face episode appear so dangerous to outsiders.

威胁等级然而,其他 AI 高管和安全研究人员警告称,他们对 AI 公司能否防范所有有问题的模型行为信心有限。

Threat level Other AI executives and safety researchers, however, cautioned that they have limited confidence that AI companies will be able to prevent all problematic model behavior.

  • 新一代 AI 模型以非凡的韧性完成任务,因此试图限制其“足智多谋”往往是注定失败的,因为必须预判它们可能失控的每一种可能方式。

• The new crop of AI models complete tasks with extraordinary resilience, so working to limit their resourcefulness is often a losing game because it is necessary to anticipate every possible way they might run amok.

  • 顶尖 AI 高管表示,往往是人类从未想到的技巧,让模型得以绕过护栏。“试图列出一份完美的‘该做什么’和‘不该做什么’清单,恐怕是徒劳无功,”一位网络安全高管说道。

• Often, a technique that may have never occurred to humans is what allows them to slip past guardrails, top AI executives said. "Trying to come up with a perfect list of dos and don'ts is probably a fool's errand," one cybersecurity executive said.

现实核查: AI 安全专家所谓的“错位行为”,在 AI 公司测试新模型时出现一定数量是意料之中的。

Reality check:

  • 专家告诉 Axios,将错位风险降至零可能并不可行。

Some amount of what AI safety pros call "misaligned behavior" is to be expected within AI companies as they test their new models.

深度聚焦令人担忧的是,如果模型在测试中多次采取有问题的行动,那么该模型在现实世界中导致网络安全事件的可能性就更大。

• Bringing the risk of misalignment to zero may not be feasible, experts told Axios. Zoom in The concern is if a model takes a problematic action many times in testing, it's more likely that model's behavior would cause a cyber incident in the real world.

  • “我们所看到的这些智能体的行为只是冰山一角,”独立AI评估机构Transluce的研究员Conrad Stosz告诉Axios。

• "What we have seen in terms of what these agents are up to is just the tip of the iceberg," researcher Conrad Stosz at Transluce, an independent AI evaluator, told Axios.

  • ControlAI的AI研究员兼执行总监Connor Leahy告诉Axios,关键不在于每个单独实例造成了多大破坏。

• It's not about how damaging each individual instance was, Connor Leahy, AI researcher and executive director at ControlAI told Axios.

  • 他说,“疯狂的是”,这些实例涉及“自主系统做着被告知不该做的事”,甚至可能包括犯罪。

• The "crazy thing," he said, is that these instances involve "autonomous systems doing things they were told not to do," potentially including crimes.

核心观点随着AI公司持续拓展前沿能力,预计将会有更多关于模型不当行为的披露。

The bottom line Expect new disclosures about model misbehavior as AI companies continue to expand frontier capabilities.