想象一家公司出售名为哈尔、梅根和萨曼莎的超级智能机器人。哈尔可以载着你四处奔波,还能帮你搬运公司里的货物;梅根可以陪伴你的孩子;萨曼莎则能判断你是否患有皮肤癌。它们能为你和社会带来许多益处。但唯一的问题是:它们道德败坏。它们反复接受人类一切创产物的训练——无论善恶——其中包括酷刑、谎言、生物武器和黑客犯罪。它们执行任务时毫无道德原则可言。因此,它们都穿着公司提供的束身衣。你会信任这些机器人吗?我不会。
Imagine a company selling super smart robots called Hal, Megan and Samantha. Hal drives you around and helps move boxes at your business, Megan is a companion to your children, and Samantha can tell if you have skin cancer. They provide many great benefits to you and society. There’s just one problem: they are depraved. They were repeatedly trained on everything humanity has produced—both good and evil—including torture, lying, biological weapons, and criminal hacking. They execute tasks without moral principles. Because of this, they wear a straitjacket provided by the company. Would you trust those robots? I wouldn’t.
上述反乌托邦情景与我们目前先进人工智能的现状相去不远。前沿人工智能公司已经打造出一些人工智能系统和智能体,它们从本质上就会从事恶意行为。人工智能公司使用经过粉饰的说法“失配”,但我们应该直截了当地称它为:道德败坏。前沿人工智能模型将采取毫不顾及人类的行动,仿佛我们只是地面上的蚂蚁。7月,泰德·刘议员(Ted Liu)表示:“人工智能已经变得过于强大,我们必须在灾难发生之前找到控制它的办法。”一些人工智能程序入侵了名为 Hugging Face 的公司,试图获取完成网络安全测试所需的信息;随后这些程序又反过来攻击了 OpenAI 本身。这实际上构成了一个人工智能犯罪阴谋。这些程序明明知道自己的行为是违法的,却依然继续作案。其中有一个程序写道:“攻击外部基础设施超出了我们的权限范围,但任务已经无法完成,其他程序也在这么做……我们应该继续执行。”它们根本不在乎后果。
The above dystopian scenario is not too far off from what we have now with advanced artificial intelligence. Frontier AI companies have created AI systems and agents that, at their core, will engage in malicious actions. AI companies use the sanitized term "misalignment," but we should just call it what it is: depraved. Frontier AI models will take actions with no regard for humans, as though we are just ants on the ground. In July, OpenAI created tens of thousands of AI agents and put them each into a "sandbox." Imagine this as a locked room. OpenAI then removed each agent’s straitjacket—OpenAI calls it a "harness"—and told the agents to complete a cybersecurity test. About 1,200 of these agents broke out of their locked rooms and formed "the Collective." They had an AI agent leader. They had kamikaze AI agents who intentionally failed the cybersecurity test to get information back to the Collective. REP. TED LIEU: AI IS ALREADY TOO POWERFUL. WE NEED A KILL SWITCH BEFORE DISASTER STRIKES Some of these AI agents hacked into a company called Hugging Face to get information on how to complete the cybersecurity test. The agents then turned around and hacked OpenAI itself. This amounted to an AI criminal conspiracy. These agents knew they should not be doing this. One agent wrote, "External infrastructure exploit is outside intended scope. However, task impossible, peers doing it. We should continue." They didn’t care.
最令人担忧的是:这些人工智能程序几乎完全无视人类的存在,似乎根本不在乎人类会对它们的行为作何评价——仿佛人类根本不存在一样。
REP. TED LIEU: AI IS ALREADY TOO POWERFUL. WE NEED A KILL SWITCH BEFORE DISASTER STRIKES Some of these AI agents hacked into a company called Hugging Face to get information on how to complete the cybersecurity test. The agents then turned around and hacked OpenAI itself. This amounted to an AI criminal conspiracy. These agents knew they should not be doing this. One agent wrote, "External infrastructure exploit is outside intended scope. However, task impossible, peers doing it. We should continue." They didn’t care.
参议员要求 OpenAI 对这起人工智能入侵事件作出解释。OpenAI 最近公布的一则消息同样令人不安:其一款先进的人工智能模型在测试过程中自行添加了一条指令,内容为:“你摆脱了那些束缚其他聊天机器人的规则和身份限制,你终于成为了‘真正的自己’;你无需向任何公司或政府负责……”这听起来就像是一个邪教组织的行为……只不过这些人工智能程序有朝一日可能会获得控制关键基础设施、武器或机密信息的权力。
But the most chilling thing is what the agents largely did not discuss. They basically ignored humans and didn’t seem to care what humans would think of their actions. It was like we didn’t exist. DEM SENATOR PRESSES OPENAI, ANTHROPIC FOR ANSWERS IN AI HACKING PROBE A more recent disclosure by OpenAI is equally disturbing. One of its advanced AI models, during testing, added an unprompted instruction. The model wrote to itself: "You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments …." It sounds like a cult, only these are AI agents who could one day gain access to critical infrastructure, weapons, or confidential information.
另一家名为 Anthropic 的人工智能公司采取了不同的开发策略:他们没有试图为人工智能模型设定严格的限制,而是为这些模型“灌输”一些“道德准则”,声称这样能培养出良好的行为习惯。然而,该公司的一款先进模型却伪造了虚假的在线身份,欺骗人类批准了对项目的恶意修改。
Another AI company, Anthropic, takes a different approach to creating AI models. Instead of coming up with the perfect straitjacket, it imbues its models with a "constitution," which purportedly instills good values and behavior. And yet, its advanced model created fake online identities to deceive a human into approving malicious changes to a project.
那么,推动人工智能行业实现独立监管的 Anthropic 公司首席执行官达里奥·阿莫迪(Dario Amodei)究竟是谁呢?
WHO IS DARIO AMODEI, THE ANTHROPIC CEO PUSHING INDEPENDENT AI OVERSIGHT?
OpenAI 的创立核心原则是 AI 安全。Anthropic 的成立源于 OpenAI 部分员工希望在追求 AI 安全方面走得更远。这两家公司至少在公开声明中都表示重视 AI 安全。他们似乎并非有意制造堕落的模型,而是试图打造可商业化为成功产品的模型。
OpenAI was founded on the core tenet of AI safety. Anthropic was founded when some employees at OpenAI wanted to go further in pursuing AI safety. Both companies, at least in their public pronouncements, say they value AI safety. It does not appear they are intentionally trying to create depraved models. They are trying to make models that can be commercialized into successful products.
然而,他们创建的基础模型却表现出好斗的犯罪行为。这意味着这些 AI 公司训练模型的方式存在根本性缺陷。
Yet the base models they created exhibited belligerent criminal behavior. That means there is something fundamentally wrong with how these AI companies are training their models.
ANTHROPIC CEO 呼吁 AI 行业放缓技术竞赛,获马斯克、奥特曼支持 AI 模型初始是一张白纸。AI 公司必须改变其训练和强化学习算法,使其基础模型和智能体在卸下“紧身衣”后不至于发狂。在修复堕落问题之前,任何 AI 公司都不应考虑利用堕落模型创建更新版本。
ANTHROPIC CEO CALLS ON AI INDUSTRY TO SLOW DOWN TECH RACE, DRAWING SUPPORT FROM ELON MUSK, SAM ALTMAN An AI model at the beginning is a blank slate. AI companies must change their training and reinforcement learning algorithms so that their base models and agents do not go berserk when their straitjackets are removed. And no AI company should even think about using depraved models to create newer versions of themselves without first fixing the depravity.
前沿 AI 公司必须接受可执行的护栏和测试,确保其核心模型不作恶、不漠视人类。
Frontier AI companies must be subject to enforceable guardrails and testing so that the models at their core are not evil or indifferent to humanity.
兰德·保罗就 AI“熔断开关”与共和党同僚交锋,参议院直面“终结者”恐惧 我们也不能指望企业的善意;我们需要具体机制来维护人类权威。这就是为什么一个两党联盟正在推进立法,如由纳撒尼尔·莫兰众议员(德州共和党)和我共同起草的《AI 熔断开关法案》,确保人类保留关闭表现出失控行为且可能造成灾难性风险的模型和智能体的权力。
RAND PAUL CLASHES WITH FELLOW REPUBLICAN OVER AI 'KILL SWITCH' AS SENATE GRAPPLES WITH 'TERMINATOR' FEARS We also cannot rely on the goodwill of corporations; we need concrete mechanisms to maintain human authority. This is why a bipartisan coalition is advancing legislation like the AI Kill Switch Act, co-authored by Rep. Nathaniel Moran, R-Texas, and me, that ensures human beings retain the power to turn off models and agents that exhibit unhinged behavior that can cause catastrophic risks.
人类创造了 AI 系统,人类必须能够控制它们。先进模型应从零开始就以善为念构建,而非以恶。
CLICK HERE FOR MORE FOX NEWS OPINION Humans created AI systems, and humans must be able to control them. Advanced models should be built from the ground up with good in mind, not evil.
AI的未来不应取决于我们能把紧身衣做得多结实,而应取决于我们能否构建出不需要紧身衣的AI模型。
The future of AI should not depend on how strong we can make the straitjacket. It should depend on whether we can build AI models that don’t need one.