
关注公众号,点击公众号主页右上角“ · · · ”,设置星标,实时关注旺材芯片最新资讯
AI机架的功率密度正在快速攀升,液冷也从“新技术”变成了AIDC真正的生命线。问题在于,当单柜功率迈向数百千瓦甚至MW级,一次水泵故障、一次过滤器堵塞、一次冷却液劣化,都不再只是普通的暖通异常,而可能在极短时间内演变成算力降频、业务中断甚至设备损坏。
这也是本文最值得关注的地方:液冷可靠性不能只盯着某一台CDU、某一只水泵或某一段管路,而要看整个系统有没有“扛故障”的能力。冗余泵能否及时接管?蓄冷能否撑过瞬态?传感器能否提前发现异常?冷却液是否真正得到全生命周期管理?
未来AIDC拼的不只是“能不能把热带走”,而是谁能在故障真正来临时,仍然把温度稳住、把算力守住。液冷的下一场竞争,本质上是一场可靠性竞争。
AI数据中心的建设成本十分高昂,单个机架的成本就可能达到300万美元甚至更高。随着未来AI数据中心逐步迈向1 GW规模,整个数据中心设施的投资预计可能达到380亿美元。面对如此巨额的基础设施投入,如何保护这些资产至关重要。因此,AI数据中心运营商正在大量部署液冷基础设施,以及时带走GPU产生的高强度热量。
AI data center deployments are expensive, withindividual racks costing $3 million or more. As AI data centers reach 1-gigawatt capacity in the near future, facilities will cost an estimated $38 billion. Protecting this kind of investment is paramount, which is why AI data center operators invest in liquid cooling infrastructure to dissipate the intense heat generated by GPUs.
对于运营商而言,最不希望发生的情况之一,就是AI数据中心液冷系统出现故障。此类故障通常由水泵异常、冷却液劣化、泄漏以及传感器问题引起。其后果可能包括设备过热、因热降频导致性能下降、业务停机、昂贵的基础设施损坏,以及因违反客户服务等级协议(SLA)而产生的赔偿或处罚,严重时甚至可能导致合同终止。如果系统发生泄漏,不仅可能损坏IT硬件;当冷却液大量泄出时,还可能需要开展环境清理工作。
The last thing an operator wants is a liquid cooling failure in AI data centers, which is typically caused by pump malfunctions, coolant degradation, leaks, and sensor issues. This can result in overheating, reduced performance due to thermal throttling, downtime, costly infrastructure damage, and customer service level agreement (SLA) penalties. Contract terminations are also a possibility. System leaks can damage hardware and may even require an environmental cleanup if enough coolant is released.
当冷却基础设施承载高密度AI计算负载时,一旦发生故障,其经济损失会进一步放大。因此,液冷系统的可靠性已经成为AI数据中心设计阶段必须重点考虑的核心问题。
When cooling infrastructure supports high-density AI workloads, the financial stakes become even higher, making liquid cooling reliability an essential design consideration.

常见液冷故障可能发生在整个冷却基础设施的任何位置,无论是设施水系统(Facility Water System,FWS)回路,还是技术冷却系统(Technology Cooling System,TCS)回路,都可能出现问题。故障原因十分多样,包括供电跌落、部件腐蚀,以及冷却液维护不当等;其中,冷却液本身也会随着运行时间增加逐渐发生性能劣化。
Common failure scenarios can occur anywhere along the infrastructure, both in the Facility Water System (FWS) and Technology Cooling System (TCS) loops. The causes may vary from power drops to component corrosion to poorly maintained coolant, which can degrade over time.
当发生泄漏时,根据泄漏位置及严重程度,冷却系统可能会触发停机,以保护相关设备。对于大多数其他故障场景,系统通常会依靠冗余水泵、冗余传感器,甚至整套冗余设备及时接管负载,从而维持正常运行。如果缺少有效的冗余措施,那么在水泵故障、过滤器堵塞、变频器(VFD)异常、控制器失效,或其他机械、电气系统部件发生问题时,系统温度可能迅速升高。在高密度AI计算环境中,温度甚至可能在数秒内就达到临界阈值,留给运维人员进行人工干预的时间窗口极其有限。
When a leak occurs, cooling systems may initiate a shutdown to protect equipment, depending on where the leak occurs and its severity. In most other failure scenarios, redundant pumps, sensors, or entire units kick in to take up the load and continue normal operations. Otherwise, temperatures will quickly rise in an emergency such as a failing pump, clogged filters, malfunctioningVFDs, controller failures, or issues with other mechanical or electrical system components. In high-density AI environments, systems reach critical temperature thresholds in seconds, leaving little margin for operator intervention.
液冷故障的常见原因
Common causes of liquid cooling failures
硬件故障——可能由水泵故障、换热器污堵、系统内部存在空气、过滤器堵塞、腐蚀、传感器失效、电能质量不佳以及系统使用不当等问题引起。例如,在系统初次冲洗过程中直接使用CDU进行冲洗,就属于不恰当的系统使用方式。
Hardware failures – These are caused by issues such as pump failures, fouled heat exchangers, air in the system, clogged filters, corrosion, sensor failures, poor power quality, and system misuse – such as using the CDU for the initial system flush.
系统泄漏——软管和接头在制造过程中存在的机械缺陷、系统设计问题以及腐蚀,都可能引发冷却液泄漏。泄漏造成的影响取决于具体位置,冷却液既可能损坏IT设备,也可能对数据中心设施本身造成破坏。
System leaks – Mechanical defects in the manufacturing of hoses and fittings, system design issues, and corrosion can cause leaks. Depending on where the leak is, the coolant can damage IT equipment and the facility itself.
冷却液恶化——系统初次冲洗不到位、缺少合理的冷却液维护计划或维护管理不当,以及系统内部存在空气等因素,都可能导致冷却液逐渐恶化,并进一步降低整个系统的换热性能,严重情况下甚至可能对芯片等硅器件造成损伤。
Coolant degradation – Poor initial system flush, lack of or mismanaged coolant maintenance plan, air in the system, degrading the coolant and the thermal performance of the system, possibly leading to damage of the silicon.
液冷故障的预防,必须从故障真正发生之前就开始。随着AI计算负载推动机架功率密度持续提升,冷却系统对于保障数据中心持续运行的重要性也越来越高。因此,运营商需要采取更加主动的可靠性管理方式,并重点围绕三个基础环节展开:规划、监控和维护。
Preventing liquid cooling failures starts long before a problem occurs. As AI workloads drive rack densities higher and cooling systems become increasingly critical to uptime, operators must take a proactive approach built on three fundamentals: planning, monitoring, and maintenance.
规划——液冷可靠性首先从系统规划开始。良好的前期规划,尤其是合理设置冗余,可以有效降低高代价故障发生的概率。在设计液冷系统时,需要提前考虑多种潜在故障场景,同时制定规范的系统初次冲洗、充液和调试方案。具体措施可能包括:采用配置内置传感器和冗余水泵的CDU,或者采用分布式冗余架构;建立合理的冷却液维护计划;同时在TCS和FWS侧配置蓄冷水箱,以帮助系统维持合适的温度水平,并在短时负荷或运行状态突变期间提供一定的缓冲和穿越能力。
Planning – This is where it all starts. Good planning that includes redundancy helps prevent costly issues. When designing the liquid cooling system, plan for various contingencies, a proper initial system flush, fill, and commissioning. This may include CDUs with onboard sensors and pump redundancy, or a distributed redundancy arrangement, a proper fluid maintenance plan, and thermal storage tanks on both the TCS and FWS to help maintain proper temperature levels and ride-through transients.
监控——对数据中心运行环境进行全天候、持续性的状态监测十分关键,其中也包括对液冷基础设施的实时监控。设备自身的告警功能,再结合DCIM和数字孪生系统提供的全局运行视图,可以帮助数据中心团队在实际故障真正发生之前识别潜在异常,并提前开展预防性维护。
Monitoring – Round-the-clock visibility into data center environments, including the liquid cooling infrastructure, is critical. Alarms within the units and overarching views from DCIM and Digital Twins alert data center teams to the need for preventative maintenance before an issue ever materializes.
主动维护——机械系统以及关键流体管网必须按照既定周期开展规范维护。这包括定期检查闭式循环系统中的冷却液状态,同时对系统中的各类部件进行检查。保持正确的冷却液配比、持续使用同一家厂商提供的PG25冷却液,并确保添加剂的用量处于合理范围都非常重要。此外,还应根据设备实际状态及时维护或更换水泵、过滤器及其他部件。
Proactive Maintenance– Mechanical systems and critical fluid networks require proper maintenance on a set schedule. This entails regularly checking the coolant in closed-loop systems and inspecting the system’s various components. Maintaining the right cooling mix, the same PG25 manufacturer, and the right amount of additives is essential, as is servicing or replacing pumps, filters, and other components as needed.
如何提高AI数据中心的冷却可靠性?
How to improve cooling reliability in AI data centers
随着AI机架功率密度和基础设施投资规模持续增长,液冷系统的可靠性已经成为保障业务连续运行、计算性能和投资收益的重要基础。尽管冷却系统故障可能造成高昂损失,但大多数故障实际上都可以通过合理的系统设计、内置冗余机制、持续监测以及主动维护加以预防。通过建立更加主动的冷却可靠性管理体系,运营商可以有效降低运行风险,并保护关键计算基础设施资产。
来源:热能工匠
专心 专业 专注


