TCP是几乎整个互联网和大多数云计算赖以运行的数据传输协议,但一位已退休的斯坦福大学教授认为,它并不适合新兴的人工智能工作负载。为解决这一问题,他正在推广一种名为Homa的新协议。斯坦福大学计算机科学荣休教授约翰·奥斯特豪特最近在AI工程师世界博览会上表示:“TCP尽管取得了许多了不起的成就,但它并不适合数据中心。”
TCP, the data transport protocol upon which pretty much the entire web and most of cloud computing is built, is ill-suited for emerging AI workloads, a retired Stanford University professor argues. To solve the problem, he's promoting a new protocol called Homa. “TCP, for all the amazing things it has done, is not a good match for datacenters,” said John Ousterhout, a professor emeritus of computer science at Stanford University, in a recent talk at the AI Engineer World’s Fair.
鉴于全球范围内对TCP的依赖之深,摆脱TCP听起来是一项极其庞大的任务。然而,奥斯特豪特表示,将Homa添加到网络中相当简单:从其GitHub源代码编译Homa,再把该模块安装到客户端和服务器的Linux内核中即可,而且无需重启。他写道:“Homa可以与TCP并行运行,因此你可以逐步将应用从TCP迁移到Homa。”运行Homa甚至会让其余TCP应用运行得更快。
Shedding TCP sounds like an immense task, given the global reliance on the protocol. But adding Homa into a network is fairly simple, Ousterhout told The. Compile Homa from its GitHub source, then install the module into the Linux kernels on the clients and servers. No reboot required. “Homa works side by side with TCP, so you can gradually move applications from TCP to Homa,” he wrote. Running Homa even makes the remaining TCP applications run faster.
旧酒换新瓶:奥斯特豪特表示,Homa是从零开始重新思考网络应如何管理流量拥塞。该协议的研究始于贝纳姆·蒙塔泽里于2019年首次发表的一篇博士论文,他现在是谷歌的资深工程师。奥斯特豪特说,如今退出教学岗位后,他把推广Homa当作了自己的“人生使命”。Homa与TCP的主要区别在于,它基于消息而非数据流。
Old wine, new skin Homa is a clean-slate rethink of how networks should manage traffic congestion, Ousterhout said. Work on the protocol began as a PhD dissertation first published in 2019 by Behnam Montazeri, now a Google staff engineer. Now that Ousterhout has retired from teaching, he has taken on the task of propagating Homa as his “life’s mission,” he said. What makes Homa different from TCP is primarily that it is message-based, rather than stream-based.
与远程过程调用(RPC)非常相似,Homa的消息长度有明确规定。与TCP不同,Homa由接收方负责拥塞控制。接收方收到的第一个数据包包含即将到达的数据量信息,因此可以明确安排数据包应在何时发送。通过这种方式,它使用最短剩余处理时间(SRPT)算法,优先处理较短的消息而非较长的消息。奥斯特豪特称,这种方法可将较短消息的延迟降低一个数量级。
Much like the remote procedure call (RPC), Homa message lengths are explicitly defined. Unlike TCP, Homa designates the receiver to manage congestion control. The first packet the receiver gets has information on how much data is incoming. It can then explicitly schedule when packets should be sent. In doing so, it prioritizes shorter messages over longer ones using a shortest-remaining-processing-time (SRPT) algorithm. This approach cuts the latency of shorter messages by an order of magnitude, Ousterhout said.
对于较短的消息,Homa延迟的第99百分位(p99)仅为92微秒,速度是TCP的13倍;相比之下,TCP的p99延迟为1.2毫秒(测试基于数据包以80%的利用率流经一个100 Gbps网络)。奥斯特豪特表示,即使对于最长的消息,Homa的性能也达到TCP的两倍。TCP在其他众多场景中也存在明显短板。目前,奥斯特豪特正在起草该协议的IETF标准化文档,同时推动将Homa上游化至Linux内核的流程。
The 99th percentile (p99) of latency for shorter messages is 92 microseconds for Homa, which is 13 times faster than the 1.2 milliseconds p99 for TCP (based on packets swimming through a 100 Gbps network at 80% utilization). Even on the longest messages, Homa is better by a factor of two, Ousterhout said. A long list of other tasks where TCP falls short Currently, Ousterhout is drafting an IETF standardization document of the protocol, as well as working through the process of upstreaming Homa into the Linux kernel.
今年3月,该协议被向后移植到了红帽企业Linux 8版和9.5版。他还正帮助大型企业评估Homa的适用性,目前正与一家大型金融服务公司合作开发原型。不过,并非所有人都愿意仅因一些延迟问题就抛弃TCP。知名网络架构师伊万·佩佩尔尼亚克于2023年发表了一篇措辞尖锐的立场论文,质疑奥斯特豪特对TCP性能的描述,并批评Homa是一个为问题寻找方案的协议。
In March, the protocol was backported to Red Hat Enterprise Linux versions 8 and 9.5. He is also helping large companies investigate Homa’s applicability – he is currently working with one large financial services company on a prototype. Not that everyone is on board with tossing TCP over a few latency issues. Prominent network architect Ivan Pepelnjak wrote a scathing position paper in 2023 about Homa, questioning Ousterhout’s performance characterizations of TCP and critiquing Homa as a solution looking for a problem.
公平地说,对陈旧迟缓的TCP感到不满的并不只是人工智能领域。在高性能数据库领域,DPDK(数据平面开发套件)正被用于绕过TCP协议栈,以实现更快的查询。存储区域网络转而采用NVMe-oF(基于网络的NVMe),通过网络结构访问固态硬盘并提高速度,其使用的传输方式包括RDMA、光纤通道和TCP。在Web领域,谷歌开发了QUIC协议,后来成为HTTP/3的基础,用以绕过TCP的队头阻塞,让浏览器能够同时下载更多资源。
To be fair, the AI community is not the only ecosystem frustrated by the dowdy TCP. In the high-performance database community, DPDK (Data Plane Development Kit) is being used to bypass the TCP stack for faster querying. Storage area networks turned to NVMe-oF (NVMe over Fabrics) to speed access to solid-state drives over network fabrics, using transports including RDMA, Fibre Channel, and TCP. For the Web, Google devised QUIC – which became the basis for HTTP/3 – to bypass TCP’s head-of-line blocking and enable browsers to download more assets simultaneously.
高频交易和多人游戏领域也都深受TCP迟缓之苦。专用RDMA网络结构以及亚马逊云服务的可扩展可靠数据报协议也应对了TCP延迟问题。机架顶端交换机在拥塞管理方面也变得更加智能,能够设置队列阈值,并通过早期拥塞通知对数据包进行标记。
Also, the high-frequency trading and multi-player gaming communities have felt the pinch of TCP sluggishness. Specialized RDMA fabrics and Amazon Web Services’ Scalable Reliable Datagram have also tackled the issue of TCP latency. Top-of-rack switches have also gotten smarter at managing congestion, setting queue thresholds and marking packets with early congestion notifications.
控制延迟确实存在。Vint Cerf和他的胡子哥们创建了TCP,让网络中桀骜不驯的信息数据包大军有了秩序,为它们提供完善的流量控制、可靠传输、连接握手和拥塞控制。拥塞控制是指避免网络路径(包括交换机和路由器)过载,而独立的流量控制机制则防止发送方压垮接收方。TCP的数据模型建立在字节流之上,也就是不区分优先级的连续数据包流。
Control lag is real Vint Cerf and his fellow beardies created TCP to bring order to the unruly hordes of message packets across networks, giving them proper flow control, guaranteed delivery, connection handshakes and congestion control. Congestion control means avoiding overload in the network path, including switches and routers, while separate flow-control mechanisms keep senders from overwhelming receivers. TCP’s data model is built on byte streams, a continuous flow of data packets with no differentiation.
消息被序列化为一个没有优先级的单一字节流。对接收方来说,较长的字节流序列与较短的字节流序列并无区别。遭遇流量洪峰的服务器可以向发送方发出警报,让其降低输入速度,但无法充分掌握还有多少流量正在涌入。发送方会根据接收方回传确认的及时程度调节输出,但必须自行猜测应在多大程度上放慢发送速度。
Messages are serialized as a single byte stream with no priority. For the receiver, longer sets of byte streams are indistinguishable from shorter ones. A server deluged with traffic can send alerts to senders to slow input, but it has limited visibility over how much traffic is still coming in. The sender itself, which regulates its output depending on the timeliness of acknowledgments sent back from the receiver, has to guess how much to slow its roll.
对于普通互联网流量或数据中心内部的大规模传输,适度增加延迟或许可以容忍。但对于延迟敏感的AI工作负载,即使是毫秒级延迟也会让人感到不适。想想GPU。推动大语言模型(LLM)发展的前沿实验室长期以来都需要顶尖的网络性能,以完成权重梯度、模型权重、KV缓存条目和检查点等任务。然而,大数据传输如今不得不日益与智能体产生的短时突发流量,以及元数据协调、缓存查询等控制任务的流量共享带宽。
For regular internet traffic or large-scale transfers within a datacenter, modest latency increases may be tolerable. But for latency-sensitive AI workloads, even delays measured in milliseconds can sting. Think of the GPUs The frontier labs driving the development of large language models (LLMs) have always required top-notch network performance for chores such as weight gradients, model weights, KV cache entries, and checkpoints. However, large data transfers must increasingly share bandwidth with short bursts of traffic from agents and control tasks such as metadata coordination and cache lookups.
Ousterhout说:“对于这些工作负载,真正重要的是延迟。”哪怕一毫秒的延迟,也会让昂贵的GPU闲置下来。Ousterhout说:“传统协议并不适合这种环境。”那么,Homa终于找到了它要解决的问题吗?还是说,它从一开始就只是超前于所处时代?至少目前,TCP仍是冠军。®
“For these workloads, what really matters is latency,” Ousterhout said. Even a millisecond of latency causes an expensive GPU to go idle. “Legacy protocols are poorly suited for this environment,” Ousterhout said. So has Homa finally found its problem to solve? Or was it just ahead of its time all along? For now anyway, TCP remains the champ.®