人工智能(AI)的使命其实很容易理解:就是利用数据和模型来优化决策流程、自动化工作流程、加速知识发现,或是创造新的服务。然而,真正实现这些目标却相当困难,需要丰富的经验才能取得成功。训练前沿的AI模型、对现有模型进行微调、处理大规模的数据推理任务,以及支持那些具备自主决策能力的AI系统,都对基础设施、业务流程以及相关人员提出了不同的要求。
AI ambition is easy to describe - using data and models to improve decisions, automate work, accelerate discovery, or create new services. Delivering those outcomes is harder and takes a significant amount of experience to be successful. Training a frontier model, fine-tuning an industry model, running high-volume inference and supporting agentic AI place different demands on infrastructure, processes and people.
AI系统的实现方式可能有多种:有些系统采用专为单一业务场景设计的“即插即用”解决方案;而有些企业则需要采取更具战略性的方法,通过构建多租户(multi-tenant)架构来充分利用规模经济,从而为各种业务需求提供支持。这就是“AI工厂”(AI Factory)的价值所在——它将AI系统的整个生命周期(从数据采集到结果输出)整合在一起,帮助企业构建出高效、经济实惠的基础设施。
Implementing AI may include a single-purpose turnkey configuration that will accommodate one line of business, or the business may demand a more strategic approach to capitalize on the economies of scale and create AI as a multi-tenant service, designed to accommodate the multitude of mainstream business requirements. That is the value of an AI factory, bringing the complete AI lifecycle together so a cost-effective infrastructure can be designed around the mission, scaled with demand, secured as needed and operated as a dependable source of intelligence.
这种基础设施能够根据业务需求进行扩展,确保系统的安全性,并作为可靠的智能信息来源为企业的运营提供支持。
The AI factory is designed to deliver the capacity, performance, utilization and business value customers expect.
一个真正的AI工厂应该具备客户所期望的所有功能:足够的处理能力、优异的性能、高效的数据利用效率,以及显著的业务价值。它是一个集成的解决方案,涵盖了数据采集、模型开发、模型训练、模型微调、数据推理以及持续优化等所有环节。其核心目标就是能够高效、重复地将数据转化为有价值的智能信息——无论这些信息最终是用于支持临床医生、工程师的工作,还是用于推动研究进展,或是服务于公共服务或企业应用。
From AI workload to business outcome An AI factory is an integrated solution for data ingestion, model development, training, fine-tuning, inference, monitoring and continuous improvement. Its purpose is to turn data into intelligence repeatedly and efficiently—whether that intelligence supports a clinician, an engineer, a researcher, a public service or an enterprise application. Achieving that goal requires more than choosing an accelerated processor. Compute must match the business model and expected workloads.
要实现这一目标,仅仅选择一款高性能的处理器是不够的;计算资源必须与企业的商业模式和预期的工作负载相匹配。网络与数据传输系统必须确保处理器能够持续获得所需的数据;存储系统必须能够处理海量数据并保证数据传输的快速性;软件系统则需负责资源的分配与任务的协调;同时,安全机制、治理规则以及多租户管理机制必须明确界定哪些用户可以使用该系统以及他们可以访问哪些数据。
Networking and data pipelines must keep accelerators supplied. Storage must support the volume and velocity of data. Software must provision resources and orchestrate jobs. Security, governance and multi-tenancy must reflect who will use the environment and what data they can access.
最后,电力供应与冷却系统也必须能够满足系统当前及未来的运行需求(包括系统密度的增加)。
Power and cooling must support the system’s density today and as it grows.
“如果没有适当的架构、治理机制以及运营方面的专业知识,企业就无法安全地利用自己的数据,也无法将人工智能(AI)融入核心业务流程,更无法将创新转化为持久的竞争优势,”HPE高级研究员、HPC与AI销售部门副总裁兼首席技术官Thierry Pienaar表示。“客户逐渐意识到:要想最大限度地实现AI的价值,他们需要一套专为AI设计、专门构建的端到端基础设施。”
“Without proper architecture, governance and operational expertise, organizations can't safely leverage their data, take AI into core processes, or turn innovation into durable competitive advantage,” says Thierry Pienaar, HPE Fellow, Vice President and CTO for HPC and AI Sales. “Customers are realizing the fact that to derive value to its utmost extent they need an end-to-end infrastructure that's purpose-designed and purpose-built for AI.”
HPE的AI解决方案套件结合了NVIDIA的加速计算技术、网络解决方案以及AI软件,与HPE的基础设施、软件和服务相结合,帮助客户部署复杂的AI系统。该套件并未强制所有客户采用统一的配置方案,而是为不同的业务目标、运营模式和需求提供了三种灵活的选择:
HPE Private Cloud AI:这是一种即开即用的AI解决方案,适用于企业内部环境,适用于需要最多256个GPU的模型训练、数据预处理(RAG)及推理等任务。
HPE AI Factory at-scale:该解决方案专为大型企业和服务提供商设计,支持大规模的AI应用场景,可管理数百到数千个GPU资源,提供集中化的控制机制、运营透明度以及多租户管理功能,确保整个AI生命周期的合规性。
Choice starts with the workload The HPE AI Factory with NVIDIA portfolio brings together NVIDIA accelerated computing, networking and AI software with HPE infrastructure, software, services and expertise at deploying complex systems. Rather than forcing every customer into a single configuration, the portfolio provides three paths for different ambitions, operating models and requirements. ● HPE Private Cloud AI, the turnkey AI factory solution, is an enterprise-ready, on-premises AI platform for running fine tuning, RAG & inferencing workload environments that need up to 256 GPUs. ● HPE AI Factory at-scale supports model builders, service providers and large enterprises that operate across many users, workloads and GPU resources (using 100s to 10s of thousands of GPUs) with centralized control, operational visibility, and multi-tenancy over the entire AI lifecycle ● HPE Sovereign AI Factory is an HPE AI factory at-scale that adds a deep level of operational control, data security and residency, sovereign management (including optional air-gapped configurations), and built-in compliance frameworks.
HPE Sovereign AI Factory:该方案在HPE AI Factory的基础上增加了更高级别的运营控制能力、数据安全保障以及数据驻留性(包括可选的物理隔离配置),同时内置了合规性框架。它特别适用于那些拥有敏感信息、且需要在法律、监管或地理范围内对数据、基础设施、模型及运营流程实施严格管控的企业。
It is designed for large enterprises, and any other organizations with sensitive information that require strict control and compliance across data, infrastructure, models and operations within defined legal, regulatory or geographic boundaries. Each option starts with the same principle: define the workloads and desired outcomes first, then select the right technologies to support them and finally identify the required resources needed to implement such an infrastructure.
每个选择都基于相同的原则:首先明确工作负载及期望的结果,然后选择合适的技术来支持这些工作负载,最后确定实施所需资源。例如,一家部署临床辅助系统的医院与一家提供 GPU 资源的服务提供商、一家负责训练视觉模型的制造商,或一家处理敏感国家数据的政府机构,在选择技术方案时会有所不同。HPE 的 AI Factory 模型为各类组织提供了实现其目标的有效途径,同时确保了系统的性能、可控性以及未来的可扩展性——HPE 与这些合作伙伴携手合作,显著提升了项目成功的几率。
A hospital deploying clinical assistants will make different choices from a service provider offering GPU capacity, a manufacturer training vision models or a government operating sensitive national workloads. The HPE AI Factory model gives each a way to build for its mission without losing sight of performance, control or future growth – and HPE partners with these organizations to dramatically increase the likelihood of success.
随着人工智能在组织中的广泛应用,相关的运营挑战也随之增加。不同团队可能需要不同的资源配置、应用程序栈、服务水平以及数据管理策略。平台团队需要监控系统的使用情况、分配计算资源、执行管理政策,并理解各项资源的消耗情况,而无需为每个工作负载单独构建独立的基础设施。在这种背景下,“从投资阶段到实际运行环境”的转化时间(即“从规划到实现”的周期)成为评估人工智能基础设施效能的重要指标:一个组织能否以低成本、高效的方式将人工智能技术从概念阶段转化为能够产生实际价值的运营环境?
Operate the environment as one system As the use of AI expands across an organization, the operational challenges mature. Multiple teams may need different resource profiles, application stacks, service levels and data boundaries. Platform teams need to see utilization, allocate capacity, apply policy and understand consumption without creating a separate infrastructure island for every workload. The management of these differences make time-to-production an increasingly useful way to think about AI infrastructure: How quickly can an organization cost effectively move from investment vision to an operational environment generating useful intelligence?
许多企业最初试图通过扩展现有的 IT 技能来应对这一挑战。然而,要自行构建一个完整的人工智能生产环境,需要具备某些特定技能(这些技能通常是大多数企业 IT 团队目前所不具备的)。如果实施不当,人工智能系统虽然可能在技术上能够运行,但在经济或运营层面仍可能面临失败;GPU 资源可能会被闲置;数据传输流程可能会成为瓶颈;冷却或电力供应的限制可能会影响系统的正常运行或扩展能力;安全政策也可能导致敏感数据或关键工作负载无法被有效处理。
Many enterprises initially try to answer that question by extending their existing IT expertise. But building a DIY production AI environment from individual components requires skill sets that many enterprise IT organizations have never needed or required at this scale. A poorly implemented AI system may technically operate while still failing economically or operationally. GPUs can sit underutilized.
不同的、难以理解的管理系统会使得人工智能工厂的基础设施难以维护和操作。因此,构建人工智能工厂本质上是一个涉及战略规划、系统集成以及运营管理的复杂问题,而不仅仅是一系列简单的硬件采购交易。Pienaar解释道:“HPE与NVIDIA合作推出的AI工厂解决方案为企业提供了多种由双方共同开发的人工智能解决方案;这些方案依托HPE的工程技术实力,能够根据企业的具体需求来设计人工智能工厂,并优化其大规模运行时的性能。”
Data pipelines can create bottlenecks. Cooling or power constraints can limit operation and/or expansion. Security policies can prevent sensitive data and workloads from being included. Separate less understood management systems can make AI factory infrastructure difficult to operate. That is why the AI factory challenge is fundamentally a strategic, systems integration and operations problem, not simply a stream of hardware purchasing transactions. "The HPE AI Factory with NVIDIA portfolio gives enterprises a range of AI solutions co-developed with NVIDIA, backed by HPE’s engineering expertise and technical capabilities to design an AI factory around their specific needs and optimize it for performance at scale."
这种灵活性非常重要——客户可以根据自身的业务目标选择合适的架构,并随着模型、用户数量以及运营需求的变化随时进行扩展。HPE AI工厂的设计旨在帮助运营商配置和管理资源、监控基础设施运行状况、追踪资源使用情况,并支持安全的多租户运营模式。这种控制能力有助于企业将资源分配与工作负载的优先级相匹配,同时让系统在不断扩展的过程中更加易于管理。
Pienaar explains. That distinction matters; customers can choose an architecture suited to their current mission and expand it as models, users and operational requirements change. The HPE AI Factory is designed to help operators provision and govern resources, observe infrastructure, track usage and support secure multi-tenant operations.
在人工智能战略中,云服务、私有环境以及混合部署模式都扮演着重要角色。对于那些对数据主权有严格要求的企业来说,选择何种技术方案取决于成本以及他们对数据控制权的需求:数据应存储在哪里?谁有权管理这些系统?适用哪些法律规范?数据如何存储?政策如何执行?不同工作负载又需要怎样的隔离措施?
That control helps customers align capacity with workload priorities while keeping the environment easier to manage as it grows. Make sovereignty a design requirement Cloud services, private environments and hybrid approaches can all play important roles in an AI strategy. For organizations with sovereignty requirements, the decision is defined by costs and the level of sovereignty and control they need: where data and models reside, who can administer the environment, which jurisdiction applies, data residency, how policies are enforced and what level of isolation various workloads require. “Sovereign AI tools from HPE and NVIDIA give an enterprise, or even a nation state, complete control over how its AI systems are built, deployed, operated and governed,” says Kaushik Shirhatti, Vice President, AI Factory at NVIDIA.
NVIDIA的人工智能工厂解决方案为企业(甚至是国家机构)提供了对其人工智能系统构建、部署、运营及管理的完全控制权。NVIDIA人工智能工厂部门副总裁Kaushik Shirhatti表示:“对于某些企业来说,这意味着必须将敏感数据保留在本国境内。”
“For some, that means keeping sensitive data in-country. For others, it means controlling who can access systems, where workloads run, how models are governed, and which local laws apply.”
“对其他人来说,这意味着需要控制谁能够访问这些系统、工作负载在何处运行、模型如何被管理,以及适用哪些本地法律。”
HPE and NVIDIA engineer for the complete outcome HPE and NVIDIA co-engineer AI factory solutions to reduce the integration work required to deploy and operate a high efficiency enterprise AI environment.
HPE与NVIDIA的工程师们共同开发了AI解决方案,以减少部署和运营高效企业级AI环境所需的集成工作。通过将NVIDIA的加速计算技术、网络解决方案及AI软件与HPE的基础设施、云服务及支持体系相结合,该联合解决方案帮助数据科学家和开发人员将更多时间用于构建和优化AI应用程序,同时让平台团队专注于维护系统的稳定运行与控制。
By combining NVIDIA accelerated computing, networking, and AI software with HPE infrastructure, cloud operations, services, and support, the joint solution helps data scientists and developers spend more time building and improving AI applications while platform teams maintain operational production and control.
NVIDIA提供了加速计算平台、网络解决方案以及NVIDIA AI Enterprise软件套件,用于支持现代AI模型的训练、微调、推理等任务;HPE则凭借其在企业系统工程、高性能计算、管理及监控软件方面的专业知识,以及在全球范围内的技术支持和服务,为这些解决方案提供了有力保障。此外,HPE还具备丰富的经验,能够满足密集计算环境对电力和冷却系统的特殊要求。
NVIDIA provides accelerated computing platforms, networking, and the NVIDIA AI Enterprise software suite to power modern training, fine-tuning, inference, and agentic workloads. HPE contributes its expertise in enterprise systems engineering, high-performance computing, management and observability software, services, global support, financing, and years of experience in the power and cooling requirements of dense computing environments.
两家公司携手合作,能够将AI计算解决方案的整体性能提升到超越单个组件所能实现的水平。他们的目标是为特定的工作负载选择合适的GPU架构和系统设计,确保加速器能够高效运行(通过高速数据传输实现最佳性能),并提供软件及运营控制所需的工具,从而实现系统的可扩展性(而无需对整个系统进行不必要的重新设计)。
Together, the companies can optimize AI computing solutions beyond any single component. The objective is to select the right GPU architecture and system design for the workload, keep accelerators productive with high-speed data movement, provide the software and operational controls teams need, and create a path to scale without unnecessarily redesigning the environment.
HPE的AI服务涵盖了从业务规划、AI策略制定、工作负载分析到系统部署、集成、维护及后续运营的全过程支持;HPE的金融服务则可以帮助客户制定合理的采购计划,并提供灵活的折旧方案及生命周期管理方案。这些能力有助于客户在综合考虑业务目标、运营模式以及系统演进速度的前提下,做出经济上合理的技术决策。
HPE AI Services support that path from business planning, AI strategy, workload characterization and facility planning through deployment, integration, support and ongoing operations. HPE Financial Services can help with purchasing, accelerated depreciation schedules and lifecycle flexibility. These capabilities help customers make economically sound technology choices in the context of the business outcome, the operating model and the pace at which the environment needs to evolve.
部署速度至关重要,但它并不是衡量成功的最终标准。客户需要考虑工作负载就绪情况、模型性能、加速器利用率、开发人员生产力、治理、可用性、经济效益以及扩展能力。这些指标将基础设施决策与组织力求实现的成果联系起来。采用搭载 NVIDIA 的 HPE AI Factory 的理由并非每个客户都需要相同的技术栈,而是每个客户都需要一个与其工作负载、数据、运营要求和目标完美契合的人工智能环境。
Deployment speed matters, but it is not the final measure of success. Customers need to consider workload readiness, model performance, accelerator utilization, developer productivity, governance, availability, economics and the ability to expand. Those measures connect the infrastructure decision to the outcomes the organization set out to achieve. The case for HPE AI Factory with NVIDIA is not that every customer needs the same stack. It is that every customer needs an AI environment intentionally matched to its workloads, data, operating requirements and goals.
通过将 NVIDIA 在加速计算领域的领导地位与 HPE 的基础设施、软件、服务和运营专长相结合,各类组织能够选择正确的路径——从而更快地将人工智能投资转化为切实的业务成果。近期部署的搭载 NVIDIA 的 HPE AI Factory 项目——加拿大的 TELUS 主权人工智能工厂以及美国犹他大学的主权人工智能工厂——正在帮助各方克服工程挑战并推动科学进步。
By combining NVIDIA’s accelerated computing leadership with HPE’s infrastructure, software, services and operating expertise, organizations can choose the right path—and move from AI investment to meaningful business outcomes faster. Recent deployments of the HPE AI Factory with NVIDIA - TELUS Sovereign AI Factory in Canada and the sovereign AI factory at the University of Utah in the US - are helping with overcoming engineering challenges and driving scientific advances.
总之,首先应确定能够从人工智能中受益的工作负载,明确相关的数据边界和驻留要求,估计在合理时间范围内的预期规模,并确定支持人工智能基础设施所需的运营模式、资源和技能。然后,与 HPE 和 NVIDIA 合作,评估哪种路径(即开即用型、大规模型或主权型)最能满足这些要求。欲了解更多信息,请访问 HPE AI Factory | 企业级人工智能基础设施 | HPE 本内容由 HPE 和 NVIDIA 赞助
In conclusion, start by identifying the workloads that would benefit from AI, define the relevant data boundaries and residency requirements, estimate the expected scale over a reasonable timeframe, and determine the operating model, resources, and skills needed to support the AI infrastructure. Then work with HPE and NVIDIA to evaluate which path—turnkey, at-scale, or sovereign—best meets those requirements. To learn more visit HPE AI Factory | AI Infrastructure for Enterprises | HPE Sponsored by HPE and NVIDIA