字号 ·· | 护眼
techcrunch

亚马逊发布自家Jev克隆版,决策模型席卷网络Amazon releases its own Jev clone as decision models flood the web

点「原文对照」整页切到原文,或双击某段只看那段的原文。

亚马逊网络服务(AWS)发布了一款受 TypeSafe 的 Jev 启发的开源决策模型,随着 AI 开发者越来越寻求比前沿大语言模型(LLM)更适合计算机自动化的智能。

Amazon Web Services released an open-source decision model inspired by TypeSafe’s Jev, with AI developers increasingly seeking intelligence that is more suited to computer automation than frontier LLMs.

亚马逊的 Strands Decider 2B 在 OpenAI 宣布类似产品的同一周发布,它是一种高速、低成本的方法,用于在预先决定的选项中进行挑选,并提供其对选择的确信程度。该模型完全开源,现已推出,且体积小到足以在本地运行。

Amazon’s Strands Decider 2B, released the same week OpenAI announced a similar offering, is a high-speed, low-cost way to sort between pre-decided options and deliver a measure of how confident it is in its choice. The model is fully open-sourced, available now, and small enough to run locally.

亚马逊杰出工程师马克·布鲁克(Marc Brooker)在看到 Jev 并尝试构建自己对这种模型的理解后,构思了这个项目。这个自制项目非常成功——它曾短暂登上该尺寸模型 Jevbench 排行榜的首位——以至于亚马逊的工程师对其进行了完善,并作为其 Strands Labs 的产品发布,该机构致力于开发用于部署 AI 代理的新工具和协议。

Amazon distinguished engineer Marc Brooker came up with the project after seeing Jev and trying to build his own take on such a model. The homebrew project was successful enough—it briefly reached the top spot on the Jevbench ranking for models of its size—that Amazon engineers cleaned it up and released it as an offering from their Strands Labs, an organization developing new tools and protocols for deploying AI agents.

布鲁克表示,对这样一种工具的需求是在与 AWS 客户的对话中产生的,这些客户的代理工作流并不总是需要功能齐全的 LLM 的能力或成本。

Brooker says the need for a tool like this emerged in conversations with AWS customers, whose agentic workflows didn’t always require the capability or cost of a fully-featured LLM all the time.

布鲁克对 TechCrunch 表示:“最初引起我对这类模型兴趣的是,它们是工作流中完美的决策者——‘根据我目前的位置,我接下来应该做什么?’”他说,它为客户提供了“一个可以采用更可靠方式构建的工作流步骤,这要归功于置信度得分、封闭的答案域,以及更低的延迟、潜在的降低的成本。”与其他决策模型一样,Strands Decider 建立在大语言模型的“躯干”之上,在此例中为 Qen3.5-2B,但它不是生成文本,而是提供经过校准的选择。

“What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step— ‘what is the next thing for me to do here, based on where I am?'”Brooker told TechCrunch. He said it offers customers “a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers, [and is] lower latency, potentially lower cost.”Like other decision models, Strands Decider is built on the “torso” of an LLM, in this case Qen3.5-2B, but instead of generating text, it delivers calibrated choices.

TypeSafe 以经济学家威廉·斯坦利·杰文斯(William Stanley Jevons)的名字将他们的模型命名为 Jev,希望借此呼应他的理论,即某种东西(如计算机智能)成本的下降实际上会增加其需求。

TypeSafe named their model Jev after the economist William Stanley Jevons, with hopes of invoking his theory that the falling cost of something—like computer intelligence—can, in fact, increase its demand.

TypeSafe 提出该理念后,研究人员已开发出数十个类似模型,这既表明了广泛的兴趣,也引发了关于其实际价值的疑问。布鲁克(Brooker)认为,挑战在于如何在优化模型快速决策能力的同时,不损害其智能水平。

The fact that dozens of similar models have been produced by researchers since TypeSafe debuted its idea shows the wide interest, but also raises the question of how valuable they can be. Brooker suggests that the challenge will be in optimizing the model’s speedy decision-making without compromising its intelligence.

他告诉 TechCrunch:“需要找到一个非常微妙的平衡点,既要提升模型在这类任务上的准确性和校准性能,又不能削弱其理解不同语言的能力,也不能损害其拥有的知识储备,正是这些特质使其具备通用性、趣味性和实用性。”

“There is a very careful balance to be found where you want to push its performance on accuracy and calibration on these kinds of tasks, without degrading its performance on understanding different languages, on having the kind of knowledge it has, which is what makes it general purpose and interesting and useful,” he told TechCrunch.

不过,他并不认为前沿实验室一定会主导这一领域,尤其是在市场规模较小的情况下,构建有趣产品的成本仅为数百或数千美元。

Still, he doesn’t necessarily expect the frontier labs to dominate the space, especially since, with smaller markets, the cost to build something interesting is in the hundreds or thousands of dollars.

TypeSafe 的高管们则表示,他们正埋头苦干,致力于改进未来的模型。

For their part, TypeSafe executives say they are keeping their heads down and improving future models.

首席执行官兼创始人迪奥戈·阿尔梅达(Diogo Almeida)告诉 TechCrunch:“我理解人们认为这是一场淘金热,但他们可能低估了让模型真正变得智能的难度。”他表示,目前尚未看到其公司面临真正的竞争。

“I get that people think it’s a gold rush, but they might be underestimating the difficulty of making the models actually smart,” CEO and founder Diogo Almeida told TechCrunch, saying that for now, he didn’t see real competition for his company emerging yet.

“目前这一批参与者更像是希望实现某种酷炫架构的机器学习人员,而非一支致力于让智能变得实用的深度投入团队。”

“The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful.”