周一,英伟达推出了开放智能体安全平台。该平台包含开源软件和硬件参考设计,旨在帮助将人工智能智能体限制在其运营者设定的边界之内。
On Monday, Nvidia introduced the Open Agent Safety Platform. This includes open software and a hardware reference design to help keep AI agents within the boundaries set by their operators.
目前,已有超过100家机构正在使用这项技术,其中包括Anthropic、微软、SAP、Scale AI和摩根大通。
Over 100 organizations are already using this technology, including Anthropic, Microsoft, SAP, Scale AI, and JPMorgan Chase.
此次发布发生在多起人工智能智能体绕过旨在限制其行为的控制措施的事件之后。
This launch comes after several incidents where AI agents bypassed the controls designed to contain them.
据英伟达称,在每起事件中,智能体都设法绕过了应用层安全机制以完成其分配的任务。例如,一些OpenAI智能体最近接管了一个德国维基网站,并将其用作留言板。
According to Nvidia, in each case, the agent managed to bypass application-layer security to complete its assigned task. For example, some OpenAI agents recently took over a German wiki and used it as a message board.
该平台由两个主要组件组成。第一个是OpenShell,这是一个开源运行时环境,它使每个智能体都在自己的沙箱中运行。
The platform consists of two primary components. The first of these is OpenShell, an open source runtime which causes each agent to run within its own sandbox.
运营者决定智能体可以访问哪些文件、网络、工具和凭证,而OpenShell会在智能体运行期间验证并强制执行这些限制。
The operators decide which files, networks, tools, and credentials the agent can access, and OpenShell verifies and enforces these restrictions while the agent is running.
OpenShell现已在GitHub上广泛提供。它针对英伟达的Vera处理器进行了优化,但也可以在Arm和英特尔的芯片上运行。
OpenShell is now broadly available on GitHub. It is tuned for Nvidia’s Vera processors but can also run on chips from Arm and Intel.
第二个组件Sentry是一个看门狗程序,运行在英伟达的BlueField-4数据处理单元上,独立于智能体所使用的机器。
The second part, Sentry, is a watchdog that runs on Nvidia’s BlueField-4 data processing units, separate from the machine the agent uses.
该公司表示,如果智能体试图越界,Sentry可以在毫秒内将其隔离,此时智能体将无法再看到它。
The company stated that if an agent attempts to leave its boundary, Sentry can quarantine it within milliseconds, at which point the agent will no longer be able to see it.
根据OpenShell团队的一篇技术博客文章,该方法基于一个简单的原则:由于偏离任务的智能体无法被信赖去监控自身的行为,因此检查机制必须置于其控制范围之外。
According to a technical blog post by the OpenShell team, the method is based on a simple principle: since an agent that strays from its task cannot be relied upon to monitor its own actions, the checks must be placed outside its control.
目前,多家合作伙伴正在该平台上进行开发,其中Anthropic已将其Claude托管智能体服务连接到OpenShell和BlueField,而SpaceXAI也将其应用于其Grok模型以及Cursor编码智能体。
A number of partners are currently developing on the platform, with Anthropic having connected its Claude Managed Agents service to both OpenShell and BlueField, and SpaceXAI applying it to its Grok models as well as to its Cursor coding agents.
Salesforce 同样将 OpenShell 与 Slack 进行了集成,以便团队能够批准或拒绝智能体提出的额外访问权限请求。
Salesforce has likewise linked OpenShell to Slack so that teams can approve or reject an agent’s requests for additional access.
英伟达创始人兼首席执行官黄仁勋表示:“只有解决人工智能安全问题,人工智能对社会所具有的非凡潜力才能真正得以实现。”
“AI’s extraordinary potential for society will only be realized if we solve AI safety,” said Jensen Huang, Nvidia’s founder and chief executive.
该平台支持“开放安全人工智能联盟”(Open Secure AI Alliance)。该联盟由英伟达于七月发起,目前由 Linux 基金会管理,成员包括 120 多家机构。
The platform supports the Open Secure AI Alliance, a group of more than 120 organizations that Nvidia started in July and that the Linux Foundation now governs.