字号 ·· | 护眼
theregister

法国开发者致力于解决机器人对图形用户界面的“失明”问题French dev aims to solve bots' blindness so they can understand GUIs

点「原文对照」整页切到原文,或双击某段只看那段的原文。

大多数大型语言模型(LLMs)在回答用户问题方面表现非常出色,但在操作 Windows 或 Linux 操作系统的桌面环境时却显得力不从心。法国人工智能模型开发者 H 在周一发布了两款专门用于处理图形用户界面(GUI)的计算机使用模型。在计算机发展的历史中,计算机使用方式主要可以分为三类:命令行界面(CLI)、应用程序编程接口(API)以及图形用户界面(GUI)。

Most LLMs are great at answering prompts, but fall short when it comes to navigating around the desktop in Windows or Linux. French AI model dev H unveiled a pair of computer use models on Monday aimed at handling graphical user interfaces (GUIs). Throughout computing history, computer use largely falls into three categories: command line interfaces (CLIs), application programming interfaces, and GUIs.

人工智能模型可以轻松地应用于前两类系统,但在操作那些更注重外观而非功能的桌面环境及应用程序时,仍然存在诸多挑战。H 公司开发的 Holo 4 系列模型旨在解决这一难题——这些模型体积小巧却功能强大,能够处理所有三种计算机使用场景,包括在图形用户界面中进行指针操作、点击、滚动以及输入等操作。Holo 4 是基于阿里巴巴的 Qwen 3.8 27B 和 Qwen 3.6 35B-A3B 模型进行训练的;通过监督学习与强化学习技术进行优化后,该模型在处理命令行界面、应用程序编程接口及图形用户界面方面的能力得到了显著提升。

AI agents can easily plug into the first two, but navigating desktop environments and applications that often prioritize form before function remains an ongoing challenge. H's Holo 4 family of models aims to address this challenge by enabling relatively small but capable models to tackle all three computer use scenarios including pointing, clicking, scrolling, and typing their way through graphical interfaces originally meant for us meatbags. Fine tuned using supervised training and reinforcement learning, Holo 4 is built atop Alibaba's Qwen 3.8 27B and Qwen 3.6 35B-A3B models. And by optimizing for CLIs, APIs, and GUIs, H claims that its models achieve far greater versatility than pure computer use models might otherwise.

H 公司声称,与纯粹的计算机使用模型相比,Holo 4 的通用性更强。此外,H 公司还更新了其基于 Nvidia Nemotron 3 架构的 Holotron 模型,该模型也具备类似的功能。在演示中,Holo 4 27B 模型成功利用 FreeCAD 的宏功能实现了埃菲尔铁塔的自动化建模(而非通过基本几何形状手动构建);而在另一个演示中,Holo 4 则通过拉伸形状来重新创建了公司的标志,展示了该模型的灵活性。

Alongside Holo 4, H has also updated its Holotron model, which is based on Nvidia's Nemotron 3, with similar capabilities. In one example, the company showed Holo 4 27B taking advantage of FreeCAD's macro function to programmatically design a 3D model of the Eiffel Tower rather than manually building it using primitives like cubes. In another demo, H did the opposite using extruded shapes to recreate the company's logo, showing the model's flexibility.

虽然这些测试结果需要进一步验证,但如果 H 公司的说法属实,那么 Holo 4 的性能确实超过了 OpenAI 等大型模型的表现——尽管其参数数量要少得多。不过值得注意的是,这并不意味着这些模型的开发成本更低;实际上,在某些情况下,这些模型的使用成本反而更高。

As with any model dev's benchmarks, take these claims with a grain of salt, but if H is to be believed, Holo 4 outperforms significantly larger frontier models from the likes of OpenAI, while using a fraction of the parameters. Curiously, this doesn't mean that they're cheaper. In fact, while the company shows higher scores, in many cases the models end up costing more per task.

根据我们对 Qwen 3.8 27B 模型的了解,其性能之所以优于 GPT 6 Luna,主要是因为 Holo 4 在生成最终结果时使用了更多的“计算资源”(即更多的“思考”过程)。虽然 GPT 6 Luna 在 OSWorld 2.0 基准测试中的表现不佳,但其训练成本要低得多。不过,这些开放源代码模型的体积非常小,因此研究人员、AI 爱好者和企业应该能够在配置相对普通的硬件上运行它们。

Given what we know about Qwen 3.8 27B, this is likely due to Holo 4 using substantially more "thinking" tokens in order to arrive at a final result relative to something like GPT 6 Luna, which doesn't perform as well in the OSWorld 2.0 benchmark, but costs substantially less. Having said that, the open weights models' diminutive size means that researchers, AI enthusiasts, and enterprises should be able to run them on relatively modest hardware.

一块 24GB 的 Nvidia RTX 3090 显卡完全足以支持这些模型以 4 位精度运行。Hugging Face 显然希望用户能够这样做:除了提供 BF16、FP8 和 NVFP4 格式的模型权重外,该公司还提供了专为 Llama.cpp(以及 LM Studio 和 Ollama)设计的 GGUF 格式模型供用户下载。模型开发者还表示,他们计划发布 DSpark 格式的模型权重,以便通过一种名为“推测性解码”(speculative decoding)的技术来加速模型推理速度。

A 24 GB Nvidia RTX 3090 should be more than capable of running these models at 4-bit precision. H clearly expects users to do just that since alongside BF16, FP8, and NVFP4 weights, it's also made a Llama.cpp (and by extension LM Studio and Ollama)-friendly GGUF version of the model available for download on Hugging Face. The model dev says that it also plans to release DSpark draft weights in order to speed up inference using a technique called speculative decoding.

我们之前曾研究过这种加速技术——简单来说,这种技术利用一个小模型来预测大模型的输出结果;如果这种技术有效,用户的模型处理和生成速度会得到提升;如果无效,则系统会自动回退到使用基础模型,从而确保输出质量不受影响。不过,这些模型本身并没有太大价值,除非有相应的运行环境或工具来配合使用。

We've explored this performance-enhancing inference tech in the past, but in a nutshell it uses a small model to guess the outputs of a larger model. When it works, users experience a speedup in token processing and generation and, when it doesn't, it falls back to the base model ensuring no loss in output quality. However, the models aren't worth much without a harness.

Hugging Face 已经开发了多个用于运行这些模型的工具,其中包含其开源的 HAI-Agents 工具包(可在 GitHub 上下载)。理论上,这些模型也应该能够与第三方开发的运行环境配合使用。Hugging Face 并不是唯一专注于计算机应用开发的模型公司;在去年的 AWS Re:Invent 大会上,该公司也发布了自己的计算机应用模型。

H has developed several agentic harnesses including its open source HAI-Agents harness, which is available for download on its GitHub. However, in theory the models should work with third-party computer use harnesses. H isn't the only model dev focused on computer use applications. At AWS' Re:Invent conference last year, the company announced its own set of computer-use models.

与此同时,美国三大模型公司——OpenAI、Google 和 Anthropic 也在这一领域进行投资,因为要突破现有的技术限制,有时只需要简单地按下某个按钮即可。

Meanwhile, the big three American model labs, OpenAI, Google, and Anthropic, are also investing in this capability, perhaps because escaping their sandbox sometimes requires pushing a button. ®