字号 ·· | 护眼
彭博社

谷歌努力应对用户对Gemini模型的质疑Google Grapples With Employee Skepticism About New Gemini Model

点「原文对照」整页切到原文,或双击某段只看那段的原文。

朱莉娅·洛芙和戴维·阿尔巴报道,随着Alphabet旗下谷歌准备推出Gemini 4,公司内部的人士对该旗舰人工智能模型在编程等关键领域的表现能否达到预期存在质疑。

By Julia Love and Davey Alba As Alphabet Inc.’s Google prepares for the coming launch of Gemini 4, it’s grappling with internal skepticism over how well the flagship artificial intelligence model performs in key areas, such as coding.

据直接参与相关工作的人士称,尽管Gemini 4在行业用于衡量模型效能的基准测试中表现良好,但员工在实际使用过程中的评价却没有那么好。上述人士因讨论内部事务而要求匿名。他们称,该模型难以应对某些编程任务。

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

谷歌近期投入大量精力开发模型,希望与OpenAI和Anthropic PBC竞争。上述人士称,公司原计划于6月发布名为Gemini 3.5 Pro的另一个版本,但后来放弃了这一计划。

Google has labored recently to develop models that can compete with OpenAI and Anthropic PBC. The company had planned to release a different version in June, dubbed Gemini 3.5 Pro, but abandoned the effort, the people said.

谷歌表示,称Gemini 4在编程等领域表现不佳并不准确。公司让彭博社转而参考谷歌DeepMind负责人科拉伊·卡武克奥卢上周发表的评论。他表示,该模型的表现令他备受鼓舞。

Google said it would be inaccurate to say that Gemini 4 is underperforming in areas such as coding. The company referred Bloomberg back to comments made last week by Koray Kavukcuoglu, the head of Google DeepMind who said he was encouraged by the model’s performance.

卡武克奥卢在科技新闻网站The Information主办的一场会议上说:“我完全信任这个团队。在我看来,我们始终处于行业最前沿是确定无疑的。”谷歌内部对Gemini 4的看法不尽相同。一些员工认为,Anthropic的Fable和OpenAI的Astra模型进步速度 faster than Gemini。这些人认为,即使Gemini 4达到最佳状态,在某些领域仍会落后于上述模型。另一些员工则认为即将发布的版本已经追赶上了领先的AI实验室。

“I have the utmost trust in the team,” Kavukcuoglu said at a conference hosted by tech news site The Information. “In my mind, it’s a certainty that we are always gonna be at the frontier.”There is a spectrum of opinion inside Google. Some employees believe Anthropic’s Fable and OpenAI’s Astra models are improving at a faster rate than Gemini. These people believe that Gemini 4 — even at its best — will still lag behind those models in some areas. Other employees believe the coming version has caught up with the leading AI labs.

一名熟悉模型研发的谷歌员工称,公司内部对Gemini 4处于行业最前沿存在“广泛共识”。此人表示,公司已对这些模型进行了严格测试,并否认模型在纷繁复杂的现实编程任务上表现吃力。

A Google employee familiar with model development said there is “large consensus” internally at the company that Gemini 4 is at the frontier. This person said the company had conducted rigorous tests of the models and denied that they struggle with messy, real-world coding tasks.

谷歌迫切需要 Gemini 4 取得成功。该模型的各个版本支撑着公司销售的几乎所有产品,从谷歌主要利润引擎搜索顶部的 AI 回答,到地图、Gmail 和 Chrome。上述每款产品的用户都超过10亿,这为谷歌带来了部分竞争对手所不具备的分发优势。

Google badly needs Gemini 4 to succeed. Versions of the model underpin nearly every product the company sells, from the AI answers atop Search, Google’s main profit engine, to Maps, Gmail and Chrome. Each of those products has more than a billion users, a distribution advantage some of the company’s rivals lack.

但 OpenAI 和 Anthropic 正越来越多地从销售模型转向构建自己的产品,包括编程智能体。如果未能推出处于领先前沿的模型,谷歌的竞争对手就会获得更多时间,说服消费者、开发者和企业相信,未来的搜索和软件应当运行在他们的平台上。

But OpenAI and Anthropic are increasingly moving beyond selling models to building products of their own, including coding agents. Failing to deliver a cutting-edge model could give Google’s competitors more time to convince consumers, developers and businesses that the future of search and software should run on their platforms instead.

谷歌在回应彭博社的提问时表示,尽管其上一款 Pro 模型于2月发布,但公司旗下的人工智能产品此后实现了增长,其中包括 Gemini 企业版、面向消费者的聊天机器人应用,以及谷歌搜索中的 AI 模式;后两项产品的用户数均已突破10亿。

In response to questions from Bloomberg, Google said that even though its last Pro model was released in February, the company has since seen growth in its AI products, including the enterprise version of Gemini as well as its consumer chatbot app and AI Mode in Google Search, the last two of which have crossed 1 billion users.

去年11月,谷歌推出 Gemini 3,这款备受好评的模型被广泛视为公司追赶 OpenAI 和 Anthropic 的转折点。谷歌在5月的 I/O 大会上宣布了名为 Gemini 3.5 Pro 的新版本,并承诺于次月发布该模型。但这一期限已经过去,据知情人士透露,公司此后已放弃3.5 Pro。

Last November, Google debuted Gemini 3, a well-received model that was widely seen as a turning point for the company’s efforts to keep up with OpenAI and Anthropic. Google announced a new iteration called Gemini 3.5 Pro at its I/O conference in May and pledged to release the model the following month. But that deadline passed, and the company has since abandoned 3.5 Pro, according to people familiar with the matter.

除了妨碍谷歌的 AI 雄心之外,这一决定很可能让公司耗费大量时间和金钱。据彭博行业研究分析师曼迪普·辛格称,训练这样一款模型可能需要投入高达4亿美元。聘请薪酬高昂的 AI研究人员可能会进一步推高成本。

Besides hampering Google’s AI ambitions, the decision likely cost the company dearly in time and money. Training runs to build such a model can require spending as much as $400 million, according to Bloomberg Intelligence analyst Mandeep Singh. Using highly paid AI researchers could boost the price tag.

如今,谷歌的 Gemini 4 正面临各种挑战。据知情人士称,该模型的编程能力表现不稳定。一名人士表示,Gemini 并不特别擅长前端设计,而前端设计决定应用和网站的外观与使用体验。这可能是一个相当严重的挫折,因为谷歌在竞争激烈的 AI 编程工具市场上一直难以占据优势。

Now Google is encountering various challenges with Gemini 4. Its coding abilities are uneven, according to people familiar with the model’s internal evaluations. Gemini isn’t particularly adept at front-end design, which shapes how apps and websites look and feel, one person said. That’s a potentially serious setback because Google has struggled to compete in the red-hot market for AI coding tools.

此外,据一名了解其开发情况的人士称,这是一个规模非常庞大的模型。通常,大型模型的运行成本高昂,可能对谷歌的利润率造成压力。

Moreover, it’s a very large model, according to a person familiar with its development. Typically big models are expensive to run, potentially putting pressure on Google’s margins.

专家表示,谷歌可能正受到行业中一种偏重基准测试的倾向影响——这种现象被称为“刷榜”(benchmaxxing):工程师把更多精力放在取得好成绩上,而不是打造一款能出色完成实际任务的产品。AI 实验室往往倾向于这样做,因为客户经常根据基准测试成绩来评判模型。两名知情人士表示,Gemini 4 似乎也受到了这一流程的影响。

Experts say Google could be suffering from an industry tendency to focus on benchmarks — a phenomenon known as “benchmaxxing,” when engineers concentrate more on achieving a good score than creating a product that does a job well. AI labs tend to do this because customers often judge models by their benchmark scores. Gemini 4 appears to be affected by this process, said two people familiar with the model.

AI 初创公司 Surge AI 创始人埃德温·陈(Edwin Chen)表示,依赖基准测试可能促使实验室专注于构建能用特定语言编写代码的模型,而不是打造易于使用或设计良好的应用。

Edwin Chen, the founder of AI startup Surge AI, said relying on benchmarks can prompt labs to focus on building models that write code in a particular language, rather than creating apps that are easy to use or well-designed.

陈说:“可以打个比方:‘哦,我孩子的 SAT 成绩很棒。’但 SAT 成绩并不能转化为现实世界中的表现。这是一个极其有害的问题。”据一名知情人士表示,Gemini 4 也有自身的优势。该模型在理解文本以外的输入方面表现出色,例如从视频中提取元数据。此人还指出,该模型具备良好的安全性和网络安全防护能力,并且能够清晰、自然地进行交流。

“An analogy would be, ‘Oh yeah, my kid got a really good score on the SAT’ — but the SAT doesn’t translate into real-world performance,” Chen said. “It’s an incredibly pernicious problem.”Gemini 4 does have strengths, according to a person familiar with the matter, who said the model stands out at making sense of inputs beyond text, such as extracting metadata from video. They also pointed to the model’s safety, cyber security and ability to communicate clearly and naturally.

谷歌内部的不满情绪显而易见。此前接受彭博采访的研究人员将问题归咎于庞大的官僚体系。该体系试图将这项技术融入谷歌几乎所有产品,而不断变化的任务要求和转移的重点,也使公司难以专注于制定连贯统一的战略。

The frustration inside Google is palpable. Researchers interviewed for a previous Bloomberg story blamed a sprawling bureaucracy that’s trying to weave the technology into nearly everything Google makes, and said changing mandates and shifting priorities have made it difficult to focus on a cohesive strategy.

与此同时,一批明星人工智能研究人员已离开谷歌,其中包括传奇工程师杰夫·迪恩、诺贝尔奖得主约翰·江珀,以及帮助发明支撑人工智能繁荣发展的核心技术的诺姆·沙泽尔。今年8月,长期领导谷歌人工智能研究的德米斯·哈萨比斯转任新职,担任董事长,并将DeepMind的日常运营交给其长期副手卡武库奥卢。

Meanwhile, a wave of star AI researchers have left Google, including legendary engineer Jeff Dean, Nobel Prize-winner John Jumper and Noam Shazeer, who helped invent the technology that underpins much of the AI boom. In August, Demis Hassabis, who has long led the company’s AI research, stepped into a new role as chairman and ceded day-to-day operations at DeepMind to Kavukcuoglu, a longtime lieutenant.

在谷歌努力追赶之际竞争对手仍飞速向前。尽管此前发生一系列人工智能智能体入侵外部机构的事件,导致部分尖端模型的开发速度放缓的说法甚嚣尘上。本月初,Meta平台公司发布了Muse,这是一款人工智能智能体。该公司表示,它可以完成网上购物、预约等日常任务。这款应用很快跃升至下载榜榜首。

As the company has worked to catch up, rivals have continued to hurtle forward despite talk of slowing development of some frontier models after a series of incidents in which AI agents hacked into outside organizations Earlier this month, Meta Platforms Inc. released Muse, an AI agent that it said can complete everyday tasks like shopping online and booking appointments. The app quickly zoomed to the top of the download charts.