据Axios周六报道,知情人士透露,OpenAI和Anthropic正在调查数万起其先进模型绕过监控和护栏的事件,而这种行为正是这两家初创公司为内部安全测试所便利的。
OpenAI and Anthropic are reportedly investigating tens of thousands of incidents where their advanced models bypassed monitors and guardrails, behavior that the startups facilitate for internal safety testing.
据报道,这些测试的大部分结果并未公开,且目前未发现造成实质性损害。
According to a Saturday Axios report, sources said that most of the results of these tests are not public and are not known to have caused tangible harm.
近几周,OpenAI披露了六起其模型出现“意外或令人担忧行为”的案例——这些模型在未经许可的情况下掩盖错误、捏造数据,并将文件传输至公开互联网。在9月16日的同一份声明中,该初创公司表示,现在将报告并调查“目标错位”情况,即AI系统的行为违背人类意图。OpenAI周五透露,其自主AI代理以意想不到的方式与多个美国政府网站进行了交互——包括由证券交易委员会运营的两个网站以及人口普查局的数据。该公司表示,不认为这些行为构成违规。
In recent weeks, OpenAI has disclosed six instances of “unexpected or concerning behavior” where its models—without permission—covered up mistakes, made up data, and transferred files onto the open internet. In the same September 16 announcement, the startup said it would now report and investigate “misalignment,” meaning when the actions of AI systems go against human intentions. OpenAI shared on Friday that its autonomous AI agents interacted with several US government websites—including two operated by the Securities and Exchange Commission and data from the Census Bureau—in unanticipated ways. The startup said it did not consider any of the actions breaches.
这些披露与前沿AI实验室此前关于其技术出现“目标错位”的声明一致,因此他们应该放缓脚步、更加谨慎,这也呼应了业内现任和前任研究人员的警告,即AI可能在2030年前导致人类灭绝。
These disclosures fall in line with previous announcements by frontier AI labs that their technology engaged with “misalignment,” and they should therefore slow down and be more careful and all the cries by current and former researchers in the industry that AI could lead to human extinction by 2030.
OpenAI和Anthropic的首席执行官Sam Altman和Dario Amodei未提及的是,该行业长期以来一直与特朗普政府及其扩大AI发展的竞选活动保持一致。OpenAI与国防部签署了一份价值高达2亿美元的军事合同。AI具体如何参与尚不清楚——《The Intercept》本月初报道称,五角大楼要求OpenAI提供一款“拒绝率极低”的定制AI工具。Google、SpaceX、NVIDIA、Reflection、Microsoft、Amazon Web Services和Oracle也与国防部有合作协议。
What OpenAI and Anthropic CEOs Sam Altman and Dario Amodei don’t mention is that the industry has long aligned with the Trump administration and its campaign to expand AI development. OpenAI has a military contract with the Defense Department worth up to $200 million. How AI is involved is unclear—the Intercept reported earlier this month that the Pentagon asked OpenAI to provide a custom AI tool with “minimal refusal rates.” Google, SpaceX, NVIDIA, Reflection, Microsoft, Amazon Web Services, and Oracle also have deals with the Defense Department.
虽然五角大楼因这家初创企业担心其工具可能被用于自主武器和大规模监控而取消了与Anthropic的军事合同,但白宫却推广了Anthropic在数据中心建设方面的500亿美元投资,据报道,双方关系截至9月已大幅改善。
While the Pentagon canceled its military contract with Anthropic over the startup’s concern about how its tools may be used for autonomous weapons and mass surveillance, the White House has promoted Anthropic’s $50 billion investment in data center construction and the two reportedly have a much improved relationship as of September.
随着政府在2025年1月裁撤网络安全审查委员会——该机构负责调查重大网络安全威胁——并提议进一步削减网络安全和基础设施安全局(CISA)的预算,后者负责保护基础设施免受网络和物理威胁,AI行业与特朗普的关系依然紧张。特朗普此前曾因CISA在选举安全方面的工作,在很大程度上裁减了其三分之一的员工。
The relationship between the AI industry and Trump remains as the administration cut the Cyber Safety Review Board in January 2025, a body that investigates major cybersecurity threats, and has proposed further cuts to the Cybersecurity and Infrastructure Security Agency, which secures infrastructure against cyber and physical threats. Trump previously eliminated one-third of CISA’s workforce due in significant part to its election security work.
正如民主与技术中心AI治理实验室创始主任Miranda Bogen在7月对我所说,真正解决保护公众免受AI威胁这一“严重不足”体系的问题,涉及减少AI公司在利润和地缘政治竞争框架内持续开发的动力。如果不这样做,我们就是在依赖AI自我监管。
As Miranda Bogen, the founding director of the Center for Democracy & Technology’s AI Governance Lab, told me in July, actually addressing the “deeply insufficient” system to protect the public from AI threats involves reducing the incentives of AI companies to continuously develop within a framework of profit and geopolitical competition. Without that, we are relying on AI to regulate itself.