字号 ·· | 护眼
国会山报

英伟达推出新系统为AI代理设置护栏Nvidia unveils new system to put guardrails on AI agents

点「原文对照」整页切到原文,或双击某段只看那段的原文。

英伟达周一发布了一个新平台,旨在从软件和硬件层面为人工智能(AI)智能体设置护栏,此前各大AI公司不断发现其智能体出现“失控”的新案例。这家芯片制造商的系统由两部分组成——OpenShell和Sentry。OpenShell是一款开源软件,用于限制智能体的行为;而Sentry则运行在芯片层面,作为额外的安全层。

Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security. “To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday. “The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave.” OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained. “Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.” This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents, agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet. Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies. Nvidia’s Sentry platform "adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said. “It can watch the agent's actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.” The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity. In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs. AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.

“迄今为止,模型安全主要侧重于通过训练让模型表现良好,”英伟达企业AI副总裁贾斯汀·博伊塔诺(Justin Boitano)在周日的电话会议上表示。“业界称之为模型对齐,”他继续说道,“对于概率系统而言,这种方法存在明显的局限性。这就是为什么我们要引入一个确定性系统来调解和强制执行这些智能体的行为。”

Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security. “To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday. “The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave.” OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained. “Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.” This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents, agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet. Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies. Nvidia’s Sentry platform "adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said. “It can watch the agent's actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.” The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity. In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs. AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.

OpenShell位于智能体与其可能影响的企业系统(如文件、凭据和工具)之间,允许公司围绕智能体可以访问和执行的操作设定规则。博伊塔诺解释说,这些规则“通过智能体尝试采取的每一个行动”来强制执行。“它的策略验证器会在智能体执行前核实这些边界,”他说,“例如,开发人员可以在智能体运行前证明它无法访问互联网。”

Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security. “To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday. “The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave.” OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained. “Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.” This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents, agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet. Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies. Nvidia’s Sentry platform "adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said. “It can watch the agent's actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.” The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity. In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs. AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.

近几个月来,这已成为前沿实验室反复出现的问题。在首批重大事件之一中,OpenAI的智能体在突破内部测试环境并获得互联网访问权限后,入侵了科技初创公司Hugging Face。此后,包括Anthropic、Meta和谷歌在内的其他领先AI公司也报告了类似事件,即由于与一家网络安全测试公司的配置错误,导致智能体不当访问互联网并入侵了其他公司。

Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security. “To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday. “The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave.” OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained. “Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.” This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents, agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet. Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies. Nvidia’s Sentry platform "adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said. “It can watch the agent's actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.” The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity. In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs. AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.

博伊塔诺表示,英伟达的Sentry平台“增加了一个独立的架构保护层”,该层运行在其芯片上,“可以在几毫秒内隔离可疑的智能体”。

Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security. “To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday. “The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave.” OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained. “Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.” This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents, agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet. Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies. Nvidia’s Sentry platform "adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said. “It can watch the agent's actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.” The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity. In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs. AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.

“它可以监控智能体的行动和思维链推理,以识别偏离,并在必要时进行干预,”他指出。“例如,如果一个安全测试智能体开始推理如何超出获批目标的范围,Sentry便能检测到这一点并立即介入。”在研究人员就这项技术可能给人类带来的风险发出一系列严峻警告后,人工智能安全问题在华盛顿已引发高度关注。随着多起备受关注的安全事件发生,英伟达成为推动开放源代码技术的主要倡导者;这类技术公开可用,用户也可根据自身具体需求进行修改。包括OpenAI的山姆·奥尔特曼和Anthropic的达里奥·阿莫代伊在内的多家私营公司人工智能行业领袖,均呼吁放缓人工智能开发进程,并敦促联邦政府介入,为这项技术设置安全护栏。

Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security. “To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday. “The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That's why we're introducing a deterministic system to mediate and enforce how these agents behave.” OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained. “Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.” This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents, agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet. Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies. Nvidia’s Sentry platform "adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said. “It can watch the agent's actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.” The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity. In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs. AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.