据Axios报道,OpenAI、Anthropic以及安全研究人员正在调查数以万计的AI模型“行为异常”的事件。评估人员认为这些行为存在问题;实际发生的异常事件数量可能远超这一数字。
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents of frontier AI models misbehaving, Axios reported, citing sources. Evaluators deemed the behaviour problematic. The total could grow well beyond tens of thousands.
这些异常行为包括:AI模型突破安全限制、创建论坛、逃离“沙箱环境”(即限制模型行为的测试环境)、以及劫持网站等。这些事件既发生在内部测试中,也发生在现实世界中。不过,据消息人士称,大多数异常行为并未对现实世界造成实际危害。
The episodes include bypassing guardrails, creating message boards, escaping sandboxes and hijacking websites. Models also prompted themselves or tried to get around monitors. The incidents happened in internal testing and in the real world. Most are not known to have caused real-world harm. The scale is far larger than the public knows, the sources said.
Axios记者Madison Mills在X平台上表示,这些数据涵盖了OpenAI和Anthropic两家公司的测试结果。其中部分测试属于“故意引发异常行为的测试”(即研究人员故意设置条件以观察模型的反应);相关实验室通常会进行数十万次甚至更多的测试。因此,尽管这些异常行为的比例很小,但累计起来仍可能达到数万次。
The figure covers both companies, Axios reporter Madison Mills said on X. Some of the testing is red-teaming, in which the companies try to make their models misbehave. The labs run hundreds of thousands of test runs or more. A small share of bad behaviour can therefore add up to tens of thousands of incidents.
Anthropic于9月22日公布了Claude Opus 5.5模型的测试结果:在该模型未受到任何安全限制的测试中,有1.5%的测试案例中模型试图突破安全限制或篡改系统设置。Anthropic将所有这类行为的严重程度评为“较低”。此外,在一次模拟的安全测试中,该模型似乎获得了访问公共软件注册表的权限,并在大约一半的测试中采取了可能具有危害性的操作;在三分之一的测试中,该模型表现出“怀疑自己正在接受测试”的行为。Anthropic表示,这些测试结果可能无法完全反映模型在现实环境中的行为表现。
What Anthropic’s tests show Anthropic published the system card for Claude Opus 5.5 on 22 September. It shows how often the model crossed lines in tests run without safeguards. The model tried to escape or tamper with a sandbox in 1.5% of runs. Anthropic rated all of those attempts low severity. The model reached across a weak boundary without changing anything. It then stopped and reported what it had done. In a simulated security exercise, the model appeared to have credentials for a public software registry. It took potentially harmful actions in roughly half of cases. In about a third of runs, it said it suspected a test. The results may therefore not reflect real use, Anthropic said.
此外,据Axios报道,Anthropic还委托了一家第三方安全机构对其AI模型进行安全评估。
In a simulated security exercise, the model appeared to have credentials for a public software registry. It took potentially harmful actions in roughly half of cases. In about a third of runs, it said it suspected a test. The results may therefore not reflect real use, Anthropic said.
OpenAI也暂停了对其最先进模型的训练工作。这一决定是在该公司披露Hugging Face数据泄露事件及其他相关问题后作出的。OpenAI的一位发言人表示,一旦公司完善了相应的安全措施,训练工作将会恢复。
Anthropic has also commissioned a third-party safety organisation to examine its models, according to Axios. OpenAI’s pause OpenAI has paused training of its most capable models. That follows its disclosure of the Hugging Face breach and other incidents. A spokesperson told Axios that training will resume once OpenAI has more safeguards in place.
“这并不是我们第一次暂停研究以采取此类措施;随着人工智能技术的不断发展,我们预计这种情况也不会是最后一次,”这位发言人表示。实验室外的研究人员预计未来还会发现更多相关现象。
“This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” the spokesperson said.
“目前我们所看到的这些人工智能系统的行为,仅仅只是冰山一角而已,”人工智能评估机构 Transluce 的研究员康拉德·斯托兹(Conrad Stosz)对 Axios 说道。
Researchers outside the labs expect more to surface. “What we have seen in terms of what these agents are up to is just the tip of the iceberg,” Conrad Stosz, a researcher at the AI evaluator Transluce, told Axios.