字号 ·· | 护眼
thenextweb

Modulate筹集2500万美元,销售能监听语音通话而不仅仅是转录文本的人工智能Modulate raises $25M to sell AI that listens to voice calls, not just transcripts

点「原文对照」整页切到原文,或双击某段只看那段的原文。

Modulate 是一家波士顿公司,致力于构建直接分析语音而非依赖文字转录的 AI 模型。本轮融资2500万美元,由 Future Ventures 领投,Hyperplane 和 Lakestar 也参与其中,使公司的累计融资额达到6000万美元。

Modulate, a Boston company that builds AI models to analyze speech directly rather than working from text transcripts, has raised $25M in a round led by Future Ventures. Hyperplane and Lakestar also took part, bringing the company’s total funding to $60M.

该公司由 Mike Pappas 和 Carter Huffman 创立。两人初次相遇是在麻省理工学院,当时 Mike 在走廊里解出了一道 Carter 正在研究的物理题。他们很快成为朋友,因为二人共同关注计算机如何影响人们建立联系的能力。

The company was founded by Mike Pappas and Carter Huffman, who first met at MIT when Mike worked out a physics problem that Carter was solving in a hallway. They soon became friends because of their shared interest in how computers affect people’s ability to connect.

“语音正在成为 AI 的主要交互界面,这也会带来一整套无法通过转录文本解决的问题。”首席执行官兼联合创始人 Carter Huffman 表示。他已从联合创始人 Mike Pappas 手中接任最高管理职位,后者现任董事长。

“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript,” said chief executive and co-founder Carter Huffman, who took over the top job from co-founder Mike Pappas, now chairman.

公司的主要产品 Velma 可以分析音频信号,识别的信息包括情绪、语气、意图、强调重点,以及声音是否由人工合成。随后,系统可以综合各种信号,识别欺诈企图、骚扰事件或客户不满等更高层次的事件,并且能在通话进行期间实时完成分析。

The main product, Velma, can analyze audio signals, picking up on emotion, tone, intent, emphasis, and whether the voice is synthetic. The various signals can then be combined to identify higher-level events such as a fraud attempt, a harassment case, or a frustrated customer, and this can be done in real time during an ongoing call.

Velma 并未采用单一大型模型,而是使用 Modulate 称为“集成听觉模型”的方法,针对每项任务选择并结合100多个规模较小、专攻特定领域的音频模型。

Instead of using one large model, Velma employs what Modulate calls an Ensemble Listening Model, which selects and combines more than 100 smaller, specialized audio models for each task.

该公司表示,这种方法的效率最高可达单一大型模型的1000倍;在识别实际问题时,Velma 的准确度是通用大型语言模型的两倍,同时误报数量降至其七分之一。这些比较均由 Modulate 自行完成。

The company states that this method is up to 1,000 times more efficient than using a single large model and that Velma is twice as accurate as general-purpose large language models at identifying actual problems while generating seven times fewer false alarms; those comparisons are made by Modulate itself.

其公开基准测试结果更容易核验。今年7月,Modulate 的转录模型在 Hugging Face 的 Open ASR Leaderboard 参评模型中名列第一;自3月以来,其深度伪造检测器一直位居 Hugging Face 的 Speech Deepfake Arena 榜首,等错误率为1.1%。批量转录每小时收费3美分。

Its public benchmark results are easier to check. In July, Modulate’s transcription models took first place out of 88 entries on Hugging Face’s Open ASR Leaderboard, and its deepfake detector has ranked first on Hugging Face’s Speech Deepfake Arena since March, with an equal error rate of 1.1%. Batch transcription costs $0.03 an hour.

深度伪造业务进入的市场中,克隆声音已成为一种成熟的诈骗工具。Modulate表示,医院正在使用其模型抵御利用深度伪造声音的来电者;其系统目前每月处理超过1000万小时音频,累计分析音频超过6亿小时。该公司还提供声音遮蔽功能,保护高风险岗位员工的安全,并可识别语音对话中对儿童实施的诱骗行为。

The deepfake work lands in a market where cloned voices have become an established tool for fraud. Modulate says hospitals use its models to defend against deepfake callers, and that its systems now process more than 10mn hours of audio a month, with more than 600mn hours analyzed in total. It also offers voice masking to protect staff in high-risk roles and detects child grooming in voice conversations.

Future Ventures联合创始人史蒂夫·朱尔维森表示,该公司已在“音频原生AI领域取得显著的技术领先”,而且需求正迅速从游戏领域扩展至AI智能体、安全和客户体验等领域。

Steve Jurvetson, co-founder of Future Ventures, said the company had “gained a significant technical lead in audio-native AI” and that demand was spreading well beyond gaming into AI agents, security, and customer experience.

这笔资金将用于研发、工程实践,以及通过推出新的SDK、API和合作伙伴集成来争取开发者,使开发语音产品的企业能够直接购买音频理解能力,而无需自行训练模型。

The money will go into research, engineering, and a push to court developers with new SDKs, APIs, and partner integrations, so that companies building voice products can buy audio understanding rather than train their own models.