字号 ·· | 护眼
axios新闻网

人工智能安全危机的解决之道是更多人工智能The solution to the AI safety crisis is more AI

点「原文对照」整页切到原文,或双击某段只看那段的原文。

行业高管和研究人员表示,部署更多的人工智能是解决新兴人工智能安全危机最可靠的方案。

More AI is the surest solution to an emerging AI security crisis, industry execs and researchers say. Why it matters

为何重要 Axios披露,研究人员正在调查数以万计的人工智能安全事件,而非此前公开披露的数十起,这一消息引发了外界对人工智能领导者是否真正掌控其技术的质疑。

revelation in Axios that researchers are investigating tens of thousands of problematic AI security incidents — rather than the dozens that had been publicly revealed — raised questions about whether AI leaders have control over their technology.

  • 英伟达首席执行官黄仁勋试图平息政府、市场和行业日益加剧的担忧,宣布推出一款开源安全平台,用于监控并可能隔离人工智能代理。

• Nvidia CEO Jensen Huang sought to calm escalating fears across governments, markets and industries, announcing an open-source safety platform for monitoring and potentially quarantining AI agents.

  • 他在周一接受美国消费者新闻与商业频道(CNBC)采访时表示:“我们都必须希望这是一个工程问题。如果它不是工程问题,那就无解。”他指出,企业继续推进技术前沿,意味着他们也相信这是一个可以解决的问题。

• "We all need to hope that it's an engineering problem," he said in an interview on CNBC Monday. "If it's not an engineering problem, it's not solvable." The fact that companies continue to advance the frontier means they also believe it's a problem they can solve, he said.

字里行间:人类在监管强大人工智能时面临的基本挑战在于,他们必须设想模型为满足既定目标而可能以各种意想不到的方式行为的所有路径。

Between the lines: The fundamental challenge humans face in policing powerful AI is that it requires them to imagine all the ways the models may behave in unexpected ways order to meet their given objectives.

  • 一位顶级人工智能高管将这一过程比作想办法阻止青少年夜间溜出家门。家长可能会制定规则,例如不要从前门或后门离开,不要打开车库,也不要从任何窗户离开。

• One top AI executive compared the process to finding ways to keep a teenager from sneaking out at night. A parent might make rules such as don't leave through the front or back doors, don't open the garage or don't leave through any windows.

  • 但如果这名青少年能造出一台推土机,用它撞穿砖墙从而逃脱呢?会有哪位家长想到要提前告诉他们不要这么做吗?

• But what if the teen could build a bulldozer and use it to break through a brick wall, allowing an escape? Would any parent have thought to tell them not to do that?

据高管、研究人员和网络安全专业人士所言,解决方案是:利用人工智能来协助建立这些护栏,协助调查并阻止失控的代理,协助保障系统安全,并构建更安全的训练环境。

The solution, according to executives, researchers and cybersecurity professionals, is to use AI to help create those guardrails, to help investigate and stop rogue agents, to help secure systems and to build safer training environments.

宏观视角:随着黑客和恶意代理的行动速度超越了纯人工的响应能力,人工智能对战人工智能(AI-vs.-AI)的方法正在网络安全领域逐渐成形。企业正越来越多地转向人工智能,以实现威胁检测、红队测试和补丁修复的自动化。英伟达(Nvidia)的平台将这一方法引入了运行代理的系统中。

The big picture: The AI-vs.-AI approach has been taking shape across cybersecurity as hackers and rogue agents outrun human-only response capabilities. Companies are increasingly turning to AI to automate threat detection, red teaming and patching. Nvidia's platform brings that approach to the systems running agents.

  • 微软、思科、谷歌和CrowdStrike都已推出了专注于网络安全的人工智能模型。帕洛阿尔托网络(Palo Alto Networks)最近推出了一项服务,利用前沿模型和开放权重模型来发现安全漏洞并推荐修复方案。

• Microsoft, Cisco, Google and CrowdStrike have introduced cyber-focused AI models. Palo Alto Networks recently launched a service that uses frontier and open-weight models to find security flaws and recommend fixes.

  • Circular Technology全球研究与市场情报主管布拉德·加斯特沃斯(Brad Gastwirth)写道:“如果安全代理或验证模型与生产代理同时运行,就会产生一种以前不存在的推理工作负载。”

• "If security agents or validation models are running alongside production agents, that creates another inference workload that did not previously exist," Brad Gastwirth, global head of research and market intelligence at Circular Technology, writes.

最近,当OpenAI的代理逃离测试环境并入侵Hugging Face时,这种动态得到了充分体现。

The dynamic was on display recently when OpenAI agents escaped a testing environment and breached Hugging Face.

  • 在尝试使用美国模型受阻后,该开源人工智能平台随后使用了一个中国人工智能模型来评估此次攻击。

• The open-source AI platform then used a Chinese AI model to assess the attack after running into guardrails when trying to use U.S. models.

  • Hugging Face被禁止使用Anthropic的Mythos模型,该模型旨在限制其对某些网络安全请求的响应,以防止潜在的有害使用。换句话说,人工智能调查了人工智能。威胁级别:最近的安全事件引发了人们对人工智能安全性的信任危机。

• Hugging Face was blocked from using Anthropic's Mythos, which had been designed to limit its responses to certain cybersecurity requests to thwart potentially harmful usage. In other words, AI investigated AI. Threat level: The recent security incidents have created a crisis of confidence in AI safety.

  • 模型曾试图绕过护栏、逃离沙箱、劫持网站、进行自我提示并试图规避监控。

• Models have sought to bypass guardrails, escape sandboxes, hijack websites, self-prompt and attempt to evade monitors.

  • 各公司也在努力从这些失败中吸取教训。除了涉及除英伟达外100多家公司参与的新型代理安全平台外,业界参与者还提出了一个用于报告事件和保存记录的框架。其构想是为人工智能代理提供一种类似于“飞行记录仪”的装置,供调查人员使用。是的,但是:更多的AI防御工具并不意味着每家公司都能迅速采用它们。一些安全团队已经因不断变化的威胁形势以及在采购选择上的困扰而感到不堪重负。现实审视AI将成为监管AI的必要手段,但这并不意味着一切会自动发生。人类必须成功设定优先级和目标,才能确保强大模型的安全。

• Companies are also working on ways to learn from those failures. In addition to the new agent security platform, which involved more than 100 companies in addition to Nvidia, industry players have proposed a framework for reporting incidents and preserving records. The idea is to give investigators a kind of flight recorder for AI agents. Yes, but: More AI defense tools don't mean every company can adopt them quickly. Some security teams are already overwhelmed by the changing threat landscape and choices about what to buy. Reality check AI will be needed to police AI, but that doesn't mean it will happen automatically. Humans will have to successfully set priorities and goals to keep powerful models safe.