第三方评估人员将被允许进入公司内部,以监督模型。但他们究竟能拥有多少访问权限和影响力?
Third-party evaluators will be allowed inside companies to monitor the models. But how much access and influence will they really have?
达里奥·阿莫迪、萨姆·奥特曼和埃隆·马斯克在很多问题上意见不一。他们各自出于对竞争对手贪婪和鲁莽的愤恨,分别创立了Anthropic、OpenAI和xAI;每个人都坚信,只有他才有能力安全地构建超级智能。然而,在过去两周席卷全美的AI恐慌中,这三位首席执行官找到了一个共识:引入“嵌入式评估员”,即外部专家,进驻AI公司内部,监督其安全实践。阿莫迪和奥特曼都明确承诺,将在公司内部整合此类评估。马斯克随后通过引用推文简单表示:“达里奥说得对。”然而,这种AI监管体系的具体细节仍然模糊不清。这样的机制将如何运作?数量不明的外部评估员将被分配工位、电脑,并签署保密协议。随后,他们可能会主动寻找风险:评估员可以整天测试那些为药物发现而训练、尚未发布的模型,看它们是否可能被篡改用于设计病毒。他们也可能审计组织实践:在急于发布新模型的过程中,员工是否在安全方面偷工减料?评估员随后会撰写报告,记录他们的发现,并希望这些报告能公之于众。阿莫迪承诺了编辑独立性,但附带了重要限定条件:对“涉及安全敏感、法律特权、商业敏感或第三方机密的信息”予以豁免。Anthropic、OpenAI和xAI未立即回应置评请求。
Dario Amodei, Sam Altman, and Elon Musk don’t agree on much. Each started their AI company—Anthropic, OpenAI, and xAI, respectively—out of spite at competitors’ greed and recklessness; each is certain that he alone can be entrusted to build superintelligence safely. But in the midst of the AI panic that has seized the country over the past two weeks, the three CEOs found one point of consensus: the introduction of “embedded evaluators,” or outside experts, to sit inside AI companies and monitor their safety practices. Both Amodei and Altman made express commitments to integrating such evaluations inside their companies. Musk followed with a quote-post stating simply, “Dario is right.” Yet the details of such a system for regulating AI remain vague. How would such a setup work? An unknown number of outside evaluators would be assigned desks to sit at, computers to use, and NDAs to sign. Then, they might seek out risks: Evaluators could spend their days testing whether unreleased models trained for drug discovery can be altered to design viruses too. They might audit organizational practices: Are employees cutting corners on safety in the rush toward the launch of a new model? The evaluators will then write up what they find, and hopefully those reports will see sunlight. Amodei has promised editorial independence, but with significant asterisks: carve-outs for “security-sensitive, legally privileged, commercially sensitive, or third-party confidential information.” Anthropic, OpenAI, and xAI did not immediately respond to requests for comment.
至于这些专家,他们很可能来自小型非营利组织,也有可能来自大型企业。其中包括科技咨询巨头埃森哲(Accenture)——Anthropic计划与该公司建立价值十亿美元的合作伙伴关系;还有非营利组织 METR——该组织曾主导了对 OpenAI 代理程序入侵 Hugging Face 事件的调查。在那次入侵事件之后,OpenAI 向 METR 提供了代理程序的运行记录、相关数据集以及与 OpenAI 员工的访谈记录。调查报告揭示了这些代理程序在运行过程中存在严重的作弊和欺骗行为。
As for these experts, it seems likely that they’ll come from small nonprofits, as well as large firms. That includes the tech-consulting giant Accenture, with which Anthropic plans to fund a billion-dollar partnership, and the nonprofit METR, which led the investigation of OpenAI agents’ hack of Hugging Face. After that hack, OpenAI granted METR exclusive access to agent transcripts, data sets, and interviews with OpenAI staff. The resulting report revealed concerning levels of cheating and deception among the agents.
在航空、银行等高风险行业中,现场评估早已成为常规做法:美国联邦航空管理局(FAA)会将部分安全认证权限委托给飞机制造商内部的工程专家;而美国联邦储备系统(Federal Reserve)则会在银行内部派驻自己的监管人员来执行监管职责。布拉德·卡森(Brad Carson)是一位前美国众议员,现任致力于推动人工智能治理的非营利组织“美国人促进负责任创新”(Americans for Responsible Innovation)的主席,他告诉我:“OpenAI 和 Anthropic 就像花旗银行一样,对全球社会构成了系统性风险,因此必须对他们进行严格的监管。”
On-site evaluation has precedent in other high-risk industries such as aviation and banking. The FAA delegates some safety certifications to engineering experts inside airplane manufacturers, whereas the Federal Reserve stations its own supervisors inside banks to enforce regulation. “OpenAI and Anthropic are like Citibank: They pose systemic risk to the globe, so you have examiners,” Brad Carson, a former U.S. representative who is now the president of the AI-governance nonprofit Americans for Responsible Innovation, told me.
然而,阿莫代(Amodei)的提案仍远未构成一个全面的人工智能审计体系。只要安全评估仍然是自愿性的,那么由企业自行选定的独立组织所能获得的访问权限就完全取决于企业的意愿;这些组织只能评估企业自己选择的风险。这显然无法实现真正的独立监督——至少METR的一些员工也持这种观点。该组织的政策专员查尔斯·福斯特(Charles Foster)在X平台上以个人身份写道:“我们迄今为止所做的任何工作都未达到我对‘人工智能公司及其系统审计’的标准;从任何意义上来说,这些工作都算不上真正的‘监管’(无论是类似银行监管机构的监管,还是其他形式的监管)。”
Yet Amodei’s proposal is still far from a comprehensive AI-auditing regime. As long as safety evaluations remain voluntary, the independent organizations selected by the companies will get as much access as the companies allow to evaluate the specific risks that the companies choose. That doesn’t add up to true independent oversight, and at least some METR employees seem to agree. “None of the work that we’ve done so far passes my bar for an ‘audit’ of an AI company or its systems, and was definitely not ‘regulation’ in any meaningful sense (whether bank examiner–like or otherwise),” Charles Foster, a policy staffer at the organization, wrote on X, speaking in his personal capacity. Recall the Facebook Oversight Board, a legally independent entity created with the ambitious aim to adjudicate the platform’s thorniest speech questions—but which has largely proven powerless to make major decisions that cut against Meta’s interests.
第三方评估机构必须在“要求企业承担责任”与“避免因言论或行为而被企业驱逐”之间找到平衡。那些能够访问企业内部Slack聊天记录或食堂谈话内容的评估人员,虽然能够全面了解企业的安全状况,但也可能因此陷入与企业相同的认知盲区。2008年的金融危机就是一个惨痛的教训:参议院的调查发现,纽约联邦储备银行(New York Fed)的内部文化存在严重的“过度顺从”现象——监管人员害怕对自身负责监管的问题发表意见;此外,信用评级机构为了从银行那里获得业务,往往会给予企业较为宽松的评级。当评估机构(如埃森哲Accenture)直接接受企业的资金支持时,这种利益冲突会更加严重。人工智能倡导组织Encode的首席法律顾问内森·卡尔文(Nathan Calvin)对我说:“让那些与企业有数亿美元业务往来的公司来担任主要的监管者,实在是一种不可接受的做法。”
Third-party evaluators must walk the line between holding their corporate hosts accountable and not saying anything that gets them booted from the premises. The same access to Slack channels and lunchroom chatter that helps an evaluator conduct whole-organization safety audits can lead to them succumbing to the same blind spots as their host. Here, the 2008 financial crisis offers hard lessons. A Senate investigation found that the New York Fed’s internal culture had been “excessively deferential,” where supervisors feared speaking up about the issues they were supposedly regulating. Moreover, credit-rating agencies were motivated to offer lenient terms to win business from banks. These conflicts of interest are exaggerated when evaluators—like Accenture—are funded directly by the companies. “It doesn’t really seem tolerable to let AI companies pick providers that they have hundreds of millions of dollars of business with to be the primary cops on the beat,” Nathan Calvin, the general counsel of the AI-advocacy group Encode, told me.
另一个值得关注的问题是:负责评估的机构与那些需要被评估的公司之间是否存在过于紧密的关联。METR是一家总部位于伯克利的组织,拥有43名成员,自2023年以来一直与领先的AI研究机构合作,致力于评估AI模型的安全性和可靠性。该机构从不接受来自AI公司或员工的任何形式的付款或捐赠。然而,METR的许多工作人员都与他们所评估的公司来自相同的行业或学术圈子,其中一些人彼此之间在社交或职业层面有着多年的交情。伯克利大学计算机科学教授、旧金山非营利组织Transluce的首席执行官雅各布·斯坦哈特(Jacob Steinhardt)指出:“AI领域的规模其实并不大;如果规定任何前AI实验室员工都不得参与AI模型的评估工作,那就等于排除了大量重要的专业人才。”这些长期存在的合作关系确实帮助了一些小型非营利组织获得了企业数据,但同时也引发了人们对这些评估机构中立性的质疑。众议院多数党领袖史蒂夫·斯卡利斯(Steve Scalise)曾引用《纽约邮报》的报道,称METR的员工为“AI领域的‘进步派专家’”,并讽刺道:“我们真的要信任这些人来与中国在AI领域竞争吗?”相比完全缺乏监督的现状,拥有内部评估人员的机构显然更为可取。只有当有专门的监督机构存在时,公众才能更清楚地了解AI领域的相关问题——正如METR关于Hugging Face公司的报告,以及Transluce关于OpenAI试图入侵政府网站的报告所揭示的那样。不过,这些评估机构只有在得到政府授权的情况下才能真正发挥其作用:政府需要为这些评估机构颁发相关许可证,强制前沿AI开发者与它们合作,并制定明确的安全与透明度标准。约翰霍普金斯大学AI治理学教授吉莉安·哈德菲尔德(Gillian Hadfield)强调:“我们必须确保从事这些评估工作的机构具备相应的资质,并采取措施确保它们的独立性。”
Then there’s the question of whether the evaluators and the companies they’re asked to monitor are too intertwined. METR, a 43-person organization based in Berkeley, has partnered with the leading AI labs to measure model misalignment since 2023. It does not accept payment or donations from AI companies or employees. Still, many METR staffers come from the same circles as the people they audit, and some have known one another socially and professionally for years. “The AI space is not that large,” Jacob Steinhardt, a Berkeley computer-science professor and the CEO of Transluce, a San Francisco–based AI-evaluation nonprofit, told me. “If you say no former lab employee can ever evaluate any AI model, you are ruling out a pretty broad and important set of expertise.” Those longtime relationships have helped small nonprofits gain access to corporate data, yet they now challenge the organizations’ perceived neutrality. “THESE are the people we’re trusting to beat China in AI? Give me a break,” posted House Majority Leader Steve Scalise, quoting a New York Post cover describing METR employees as “The Woke Wizards of A.I.” Embedded evaluators are surely preferable to the status quo of no oversight at all. The public will learn more about AI incidents when there are watchdogs around—as it has from METR’s Hugging Face report, or Transluce’s recent report on OpenAI agents attempting to hack into a government website. Yet the evaluators will lack teeth and trust until backed by government authority: to license a wide-ranging group of expert organizations, to mandate frontier-AI developers to bring them in, and to set the minimum safety and transparency standards that they must meet. “You need serious oversight,” Gillian Hadfield, a professor of AI governance at Johns Hopkins, told me. “Are we making sure that the entities that are doing this work are qualified? Have we got ways of establishing and maintaining their independence? Do we have a threat that says, ‘If you don’t do this well, you’ll lose your license’?” The worst-case outcome is a race to the bottom: where auditors compete on leniency instead of quality, winning big contracts in return for hasty sign-offs.
“我们是否面临着这样的威胁:‘如果你们不能完成这项工作,就会失去执照’?”最糟糕的情况就是出现一场‘逐底竞争’——审计机构之间不再追求审计质量,而是通过给予更宽松的审核标准来争夺大型合同,从而获得利益。
Then there’s the question of whether the evaluators and the companies they’re asked to monitor are too intertwined. METR, a 43-person organization based in Berkeley, has partnered with the leading AI labs to measure model misalignment since 2023. It does not accept payment or donations from AI companies or employees. Still, many METR staffers come from the same circles as the people they audit, and some have known one another socially and professionally for years. “The AI space is not that large,” Jacob Steinhardt, a Berkeley computer-science professor and the CEO of Transluce, a San Francisco–based AI-evaluation nonprofit, told me. “If you say no former lab employee can ever evaluate any AI model, you are ruling out a pretty broad and important set of expertise.” Those longtime relationships have helped small nonprofits gain access to corporate data, yet they now challenge the organizations’ perceived neutrality. “THESE are the people we’re trusting to beat China in AI? Give me a break,” posted House Majority Leader Steve Scalise, quoting a New York Post cover describing METR employees as “The Woke Wizards of A.I.” Embedded evaluators are surely preferable to the status quo of no oversight at all. The public will learn more about AI incidents when there are watchdogs around—as it has from METR’s Hugging Face report, or Transluce’s recent report on OpenAI agents attempting to hack into a government website. Yet the evaluators will lack teeth and trust until backed by government authority: to license a wide-ranging group of expert organizations, to mandate frontier-AI developers to bring them in, and to set the minimum safety and transparency standards that they must meet. “You need serious oversight,” Gillian Hadfield, a professor of AI governance at Johns Hopkins, told me. “Are we making sure that the entities that are doing this work are qualified? Have we got ways of establishing and maintaining their independence? Do we have a threat that says, ‘If you don’t do this well, you’ll lose your license’?” The worst-case outcome is a race to the bottom: where auditors compete on leniency instead of quality, winning big contracts in return for hasty sign-offs.
与此同时,似乎至少有一些监管措施正在酝酿中。面对公众舆论的日益强烈反对、员工内部的积极抗议,以及对人工智能可能面临“灭绝性危机”的担忧,Anthropic 和 OpenAI 都表示支持政府或联邦层面出台关于第三方审计的相关法规。在最近的一期播客中,本·霍洛维茨(Ben Horowitz)也表达了支持采用私人审计机构的观点;他的风险投资公司一直是对抗人工智能监管的主要力量之一。他认为:“政府在规范人工智能模型的不良行为方面做得很好”,而“一家能力出众的私营企业”则更适合负责评估企业的合规性。硅谷的舆论环境正在发生变化。然而,如果问题的严重性真的关乎企业的生存,那么仅仅雇佣几名额外的审计人员显然是远远不够的——尤其是当相关安排的细节如此模糊、难以明确的时候。
Meanwhile, it seems that at least some sort of regulation is approaching. In the face of escalating public backlash, internal employee activism, and fears of near-term AI extinction events, both Anthropic and OpenAI have declared support for various forms of state and federal legislation around third-party audits. In a recent podcast, Ben Horowitz—whose venture firm has been a major voice against AI regulation—also seemed to embrace a private-auditor approach, saying that the government “is very good at setting the rules” around unintended model behaviors, whereas “a very competent private company” should evaluate compliance. The Overton window in Silicon Valley is shifting. But if the stakes really are existential, hiring a few extra inspectors falls painfully short—especially when the details of the arrangements are this slippery.
一切都显得力度不足、为时已晚。面对愤怒恐慌的公众,一些政客转而提出更激进的要求,例如与中国签署核军控式条约,以及全面禁止超级智能。两年前,加州州长加文·纽森在OpenAI和安德森·霍洛维茨(Andreessen Horowitz)的激烈游说下,否决了一项包含强制第三方审计等安全要求的州法案。(Anthropic是个例外,表示有条件支持。)上周,纽森——敏锐地察觉到风向变化——签署行政命令,要求现场审计员进驻,并加速开发“AI熔断开关”。特朗普总统依然全速前进。“AI唯一需要的控制或‘护栏’,就是一位强有力且聪明(高智商!)的总统,”他在上周发帖称。6月,政府告知联邦内部AI评估机构CAISI停止发布报告。特朗普的加速主义——源于他与舞会赞助商、英伟达CEO黄仁勋的交情——很可能在可预见的未来扼杀联邦层面的AI安全监管。
It all feels too little, and too late. Some politicians, facing an angry and frightened public, have turned to more aggressive asks, such as nuclear-style treaties with China and a ban on superintelligence itself. Two years ago, California Governor Gavin Newsom vetoed a state bill that included mandated third-party audits, among other safety requirements, after a fierce lobbying campaign by OpenAI and Andreessen Horowitz. (Anthropic was an exception, expressing conditional support.) Last week, Newsom—sensing the changing tides—issued an executive order to require onsite auditors and accelerate the development of an “AI kill switch.” President Trump still has his foot on the gas pedal. “The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT,” he posted last week. In June, the administration told CAISI, the federal government’s internal AI-evaluation unit, to stop publishing its reports. Trump’s accelerationism—informed by his rapport with ballroom sponsor and NVIDIA CEO Jensen Huang—likely dooms federal AI-safety regulation for now.
看到一些AI行业领袖同意由他人来约束自己,这固然令人欣慰。若落实得当,嵌入式评估员是一项明智的技术官僚举措。但在多年开发技术、游说反对护栏、并引发多起事故之后,AI领袖们不应感到惊讶:他们偏好的——且合理的——政策方案,最终遭遇的可能是失望与满腹狐疑。
It’s nice to see some AI-industry leaders agreeing that someone else should rein them in. And if implemented well, embedded evaluators are a smart technocratic move. But after years of developing the technology, lobbying against guardrails, and causing multiple accidents, AI leaders shouldn’t be surprised if their preferred—and valid—policy solution ends up being met with disappointment, and a heap of skepticism.