Chapter 5.11
In this chapter · 7 sections
Thermal Design, Reliability, Leak Detection & Commissioning
A liquid-cooled hall survives only if its thermal budget closes at the worst branch and its leak and loss-of-flow response preserves thermal and containment limits through each required service or fault state; prove the local timing before the GPUs find the gap, then hand that evidence to Chapter 13.5 for installed acceptance.
What you'll decide here
- Where in the chain you spend your thermal margin — at the cold plate, the CDU approach, or the facility loop — because a budget that closes at the average branch but not the worst-case branch is a budget that throttles GPUs in production.
- How much cooling redundancy and ride-through to commission — N, N+1 or 2N pumps, UPS-backed pumping, and a participating chilled-water thermal flywheel or none — to keep the declared load running, reduce it or stop it safely before the named equipment’s thermal margin expires.
- The leak-detection and containment architecture — sensing modality, zoning granularity, and the auto-isolation response — and whether a detected leak pauses a rack, a row, or the job.
- The serviceability contract: dripless quick-disconnects, hand- or blind-mate trays, and the restore-time target that decides whether a failed cold plate is an isolated tray swap or a branch drain-and-fill; demonstrate isolation, handling, spill, air inclusion, leak checks and return to operation.
- The commissioning evidence you fund from L1 factory checks through L5 integrated testing, handed to Chapter 13.5, and the fluid-analysis cadence in Chapter 5.7 — the difference between catching a poor fill or failed response before release and discovering fouling, corrosion and derating in operation.
Everything earlier in Part 5 chose a cooling technology and sized its pieces. This chapter is where those pieces are made to work together, reliably, and survive the day something fails. It is the engineering that separates a rack that runs at rated TDP for five years from one that quietly throttles every afternoon because the thermal budget never actually closed at the worst rack in the row. The discipline here is unforgiving in a way air cooling never was: when you put water (or PG25) inside the rack at 200+ kW, a single open budget, a single undetected leak, or a single un-rehearsed loss-of-cooling event does not degrade gracefully — it can trigger a rack protection response whose timing must be established from OEM transient data, controls tests or a professional-engineering calculation and, on a synchronous job, restarts thousands of GPUs from the last checkpoint.
Four problems follow, each constraining the next. The end-to-end thermal budget comes first — a worked ladder on a named rack that walks the temperature from junction to wet-bulb with the resistance stack 5.1 defines, and proves the loop closes at the worst-case branch, not the average. Then the reliability architecture: redundancy topology, UPS-backed pumps, and the selected rack's validated thermal ride-through. Then leak detection and containment — sensing, zoning, and the auto-isolation response. Finally serviceability and commissioning — quick-disconnects, blind-mate, MTTR, the handoff to integrated acceptance in 13.5, and the fluid-analysis program that keeps the loop in spec for life. The cooling-controls transient problem — anti-hunting, setpoint stability, dew-point excursions under a synchronized load slam — is the thermal twin of the electrical transient and lives in its own chapter (Chapter 5.12); the consolidated failure-mode catalog lives in Appendix F.
The end-to-end thermal budget: closing the loop at the worst branch
The single most useful artifact in a liquid-cooled design is a temperature ladder that starts at the silicon junction and ends at the outdoor wet-bulb, carrying every thermal resistance and every approach temperature. Close that ladder — leave the junction under its throttle limit when the outdoor air sits at design wet-bulb and the worst rack in the hall runs full load — and the facility works. Close it on a spreadsheet for the average rack but not the rack at the end of the longest manifold run, and you throttle, chasing a designed-in goodput leak through the first year of operations.
Follow the heat forward. A Blackwell-class GPU dissipating on the order of 1–1.4 kW presents a die-average heat flux on the order of ~90 W/cm² (1.4 kW over roughly 16 cm² of silicon), with local hotspots well beyond that — past air’s practical socket-level limit, which is the whole reason we are here. The junction-to-coolant path runs through the silicon, the TIM, the cold-plate base, and the convective film inside the cold-plate microchannels. Each is a thermal resistance in K/W; the sum, multiplied by the chip power, is the rise from coolant to junction. Then the coolant itself heats as it crosses the cold plate — that is your cold-plate delta-T, governed by flow rate and fluid heat capacity. The warmed coolant returns through the rack manifold (a pressure-drop and mixing penalty), crosses the CDU heat exchanger (an approach-temperature penalty of typically 3–5 °C between the technology-cooling loop and the facility water loop), travels the facility loop to the heat-rejection plant, and finally rejects to ambient across a dry cooler or tower (another approach to dry-bulb or wet-bulb). Each handoff costs you degrees. The budget is the bookkeeping of those degrees against the one number that matters: the GPU’s throttle threshold. The ladder below does that bookkeeping for a named rack — Lenovo’s GB300 NVL72, whose published flow table fixes the coolant rise — with every other operand declared as an assumption; Chapter 5.1 owns the resistance stack itself, and what this chapter adds is the local limit the protection schedule must hold.
| Stage | What it is | Typical rise / approach | Running temperature |
|---|---|---|---|
| Outdoor design wet-bulb | The climate floor you reject against | — | ~24 °C (assumed design WB) |
| Heat-rejection approach | Tower/dry-cooler approach to WB/DB | +4–8 K | ~30 °C facility supply |
| Facility loop + CDU HX approach | FWS→TCS handoff across the CDU exchanger | +3–5 K | 35 °C TCS supply — Lenovo’s 35 °C row: 89 L/min per rack |
| Manifold + worst-branch penalty | Manifold heat gain and blending on the worst branch (under-flow there widens its own rise as well) | +1–3 K | ~37 °C at the worst cold-plate inlet |
| Coolant rise across the rack | 121.5 kW to liquid ÷ (89 L/min × 1.00 kg/L × 4.18 kJ/kg·K ÷ 60) = 121.5 ÷ 6.20 kW/K | +~20 K | ~55 °C at the rack return; ~57 °C leaving the last plate on the worst branch |
| Junction-to-coolant resistance | Silicon + TIM + plate + film × chip power, referenced to the coolant leaving that plate | +25–35 K (assumed) | ~82–92 °C junction |
| Margin to throttle | Headroom before clock/voltage de-rate | assumed throttle threshold 85 °C | +3 K at the low end of the assumed resistance, −7 K at the high end — the stack closes only if the cold-plate resistance, the CDU approach and the branch balance all land at their good ends; redesign any stack that cannot show positive worst-case margin |
The last two rows hold the money. Only the margin to throttle on the worst branch at design wet-bulb counts. An inexperienced designer chases a colder plant — drop the facility supply, buy back degrees with a chiller. A colder plant buys margin at an energy/capital cost; a warmer supported point can expand economizer hours. Select the operating point from the named TCS limits, FWS class, climate bins, heat-exchanger and rejection approaches, load, control range and redundancy — not from a universal W40/W45 default. Recover margin upstream of the plant instead: a tighter CDU approach, a better TIM, a higher per-chip flow to shrink the cold-plate delta-T, or balancing the manifold so the worst branch is not 3 °C hotter than the best. Spend the degree where it is cheap. → fundamentals and the metric definitions are in Chapter 5.1; the cold-plate and flow design in Chapter 5.4; the warm-water loop in Chapter 5.7; heat rejection in Chapter 5.8.
Reliability architecture: redundancy, UPS-backed pumps, and the ride-through problem
A direct-to-chip loop's loss-of-flow response is product- and implementation-specific. Establish whether the selected rack throttles, performs a controlled or emergency shutdown, or trips—and on what timescale—from OEM transient data, controls tests or a professional-engineering calculation. That evidence sets the pump and heat-rejection protection and redundancy posture; no universal seconds-to-trip value or N+1/2N topology is implied.
The redundancy question then splits into two layers that fail on completely different timescales. The pumping layer (CDU pumps, secondary-loop circulation) must be sized and protected against the selected rack's validated loss-of-flow response; whether pumps need UPS/BESS, protected feeds or redundant circuits comes from that evidence. The heat-rejection layer (chillers, dry coolers, towers, the facility loop) has a product- and implementation-specific response because loop volume and temperature margin vary; validate that transient rather than assuming a portable ride-through time. The design choice is whether to buy back ride-through with thermal storage — a buffer tank, a chilled-water volume, a deliberately oversized loop — or to accept the thin margin and lean on fast, reliable failover. → the electrical spine that powers all of this, and the BBU/BESS layering that backs it, is in Chapter 4.5; CDU sizing and N+1 pump topology in Chapter 5.6; heat-rejection redundancy and the lack of chilled-water inertia in DLC in Chapter 5.8.
| Posture | Pumping (CDU) | Heat rejection | Ride-through | Selection condition |
|---|---|---|---|---|
| N | Single pump path | No spare plant | Test-specific response; define trip/shutdown state | Only where every permitted loss and recovery remains inside the thermal and service objective |
| N+1 | Redundant pump in CDU, shared feed | One spare chiller/cooler | Test-specific response | Where one named pump, CDU, or plant-unit state must retain required flow or capacity |
| N+1 + UPS pumps | Redundant pumps, dual UPS-backed feeds | N+1 plant | Pumps ride the UPS; plant fails over | Where source-transfer interruption exceeds the thermal ride-through without backed pumping |
| 2N + thermal store | Fully duplicated pump paths | 2N plant + buffer-tank flywheel | Stored cooling sized from the tested transient | Where an independently defined full-path loss must retain contracted thermal service |
The cooling posture is a state-based continuity decision. Checkpointing can reduce the goodput consequence of a trip, while request retry and fleet failover can reduce serving impact; neither proves that N, N+1, or 2N cooling is sufficient. Define the maintenance and fault states, the tested throttle/shutdown response, transfer interruption, post-event flow and rejection capacity, shared power and controls, recovery, and contract. Compare the incremental topology cost with workload and fleet recovery in Chapter 12.5.
Leak detection and containment: the cascade you must never start
Water and energized electronics are an old enemy, and the industry's tolerance for leaks in a 200 kW rack is effectively zero. The containment philosophy is layered defense: keep the fluid in the loop; if it escapes, detect it within the validated local window and coordinate heat reduction, containment and surge-approved isolation so the released coolant cannot reach a live busbar; and at every stage, make the response proportional so a drip pauses a rack rather than a job. The connectors do most of the work — modern in-rack plumbing is built on dripless quick-disconnects (UQD/UQDB-class couplings) and blind-mate floating-tray connections that self-align and seal on insertion, so that the act of servicing a node does not itself create a spill. But connectors fail, hoses chafe, cold plates corrode, and gaskets age, so detection and isolation are the backstop.
Detection runs on three complementary modalities, and serious designs use more than one. Point sensors — float switches and conductive pads in drip trays and at low points — are cheap and certain but only see fluid that has already pooled where you guessed it would. Rope/cable leak-detection snakes along manifolds, under racks, and through cable trays, localizing a leak to a length of cable rather than a point. Pressure and flow telemetry on the loop catches the leak that never reaches a sensor at all — a slow drop in loop pressure or a mismatch between supply and return flow betrays a leak upstream of any pooling. The 2026 frontier adds ML/IoT leak forecasting: correlating pressure, flow, makeup-water consumption, and acoustic signatures to flag a degrading connector days before it weeps. The response must be zoned: the selected local response should alarm and isolate the affected rack or row by closing its manifold valves in the approved heat-removal and surge sequence — not silently dump the whole hall's cooling, which would convert a contained drip into a hall-wide loss-of-cooling trip.
Deep dive: the coolant-leak cascade and the thermal-runaway twin (the two failure modes that define DLC risk)
Two cascades dominate the DLC failure catalog, and they are mirror images. The coolant-leak cascade starts with a breach — a chafed hose, a failed gasket, a cross-threaded quick-disconnect, a corroded cold-plate channel. Fluid escapes onto energized electronics; if PG25 reaches a busbar or PSU it can flash a short, and if loop pressure drops far enough the remaining racks on that branch lose flow. The contained version stops at the drip tray with a rack isolated and a service ticket. The uncontained version is a short, a fire-risk event, and a multi-rack loss-of-cooling trip in the same minute. The entire detection-and-isolation architecture above exists to keep this cascade in its contained form. The insurability and FM Global gating that shapes fluid choice (and that stalled two-phase immersion) is the financial expression of this same risk — see Chapter 5.5.
The thermal-runaway cascade is the leak's twin: instead of fluid leaving the loop, heat fails to leave the silicon. A blocked filter, a fouled cold plate, a seized pump, a closed valve, or a controls hang stops or starves flow with the loop still full. The junction response after flow loss is product- and implementation-specific; establish whether it throttles, shuts down or trips from OEM data, controls tests or an engineering calculation. On a synchronous job, one rack's trip cascades into a job-wide checkpoint restart across thousands of GPUs. The defenses differ from the leak case — they are flow/pressure health telemetry, redundant pumping, fast power-capping, and thermal storage — but the lesson is the same: the failure is fast, so the detection and response must be faster than the thermal time constant. Both cascades, treated as explicitly dual-use (random fault and attacker-induced — a maliciously closed valve looks identical to a seized one), are consolidated in the FMEA catalog in Appendix F. The transient-controls failure that can cause a thermal-runaway trip — a hunting valve, a setpoint oscillation — is in Chapter 5.12.
Serviceability and MTTR: the hot-swap contract
A liquid-cooled rack will need service — a cold plate fouls, a hose ages out, a node fails — and what decides your fleet availability is how long that service takes. Mean-time-to-repair is a design parameter you set at procurement, not a number you discover in operations. The two architectures that bracket it: a rack built on dripless quick-disconnects and blind-mate trays lets a technician pull a node, swap the failed assembly and re-seat it through a defined tray procedure whose isolation, tools, spill, air inclusion and restoration time have been demonstrated — the connectors break and seal cleanly, and the rest of the rack keeps running. A rack plumbed with conventional fittings and shared manifold segments that cannot isolate forces a drain-and-fill of a whole branch to touch one node — hours of downtime, a re-bleed to purge air, and a re-commission of that branch before it carries load again.
The serviceability contract is therefore a set of concrete commitments to extract at procurement: every fluid connection is a dripless quick-disconnect rated for the cycle count of expected service; each node uses its qualified hand-mate or blind-mate arrangement — where blind-mate is selected, insertion seats the intended power and coolant interfaces within their alignment limits; every branch has isolation valves so a single cold plate can be isolated without draining the row; and the CDU carries side-stream filtration (Deschutes-class designs run 0.2-micron filtration) and redundant pumps so filter and pump service never stop cooling. OCP is standardizing exactly these interfaces — rack/manifold geometry, UQD/UQDB couplings, leak-detection practice — precisely so a hall is not locked to one vendor's serviceability story. The cross-vendor interoperability of couplings, manifolds, and coolant chemistry remains an open standardization gap and a real fragmentation risk to weigh in procurement.
Scope & caveats
Vendor-reported availability of Google's own liquid-cooled TPU CDU fleet (Project Deschutes architecture, redundant pump and heat exchanger, UPS-backed) since 2020, across more than 2,000 TPU Pods. Google states neither the availability definition, the excluded-event list nor the measurement method, and the figure is not independently audited; it is not a rating for a purchased CDU and does not transfer to a non-redundant or non-UPS-backed installation.
Scope & caveats
Published per-rack flow requirement at the named supply temperatures; the ~20 K rise the 35 °C row implies for the liquid share of 135 kW is the guide’s arithmetic, not a Lenovo figure.
Scope & caveats
Reported forecast estimate, not a measured deployment census or a project cooling-selection rule. The cited PMR cold-plate category is broader than single-phase DTC.
Reported forecast estimate, not measured fleet share; PMR's published cold-plate category is broader than single-phase DTC and is not a project-selection rule.
Local evidence handed to integrated acceptance
Commissioning is where the design meets reality: finding a weeping weld before delivery is cheaper than finding it over a live rack. Hand the local cause-and-effect schedule to Chapter 13.5 for integrated cooling acceptance; 13.6 owns the L1–L5 ladder. Before local response tests, close the isolated pressure-system boundary, flush out manufacturing debris to the fluid-release limits, calibrate sensors and synchronize timestamps. Establish commanded heat and flow at the worst branch, then use approved signal injection or a qualified thermal simulator to exercise pump loss, a leak signal or a control fault. An artificial heater proves its heat-transfer and control response, not an accelerator’s no-flow junction transient.
Distinguish command from physical result: record detection, actual heat reduction, valve position, surviving branch flow, temperature, pressure and reset permission on one time basis. Sensor uncertainty and timing resolution spend margin. A delivered alarm with no required heat or valve response fails, even if the test happened to finish before anything overheated.
Does the local leak response meet its stated limits?
Heat-off upper bound = 1.80 + 0.05 = 1.85 s, inside 2.00 s. Flow lower bound = 112 × 0.980 = 109.76 L/min, above 109 L/min. Temperature upper bound = 36.2 + 0.2 = 36.4 °C, below 38.0 °C. Pressure bounds are 0.09–0.51 MPa(g), inside the stated range. Retain the sequence as a candidate for test; these assumed values do not certify hardware survival.
Repeat the proposed case with the supervisory connection absent and verify that actual local heat-off, isolation and deliberate reset meet the same limits. Capture command and actual heater power, valve position and alarm state alongside the thermal signals. A 10 Hz thermal log can support this example’s stated event resolution; select a separate pressure sampling rate from the shortest relevant hydraulic wave transit. A slow thermal trace cannot clear surge.
The heat-off crossover is an indicated 1.95 s when ±0.05 s timing uncertainty applies. A failing illustrative trace at 2.40 s has an upper bound of 2.45 s and fails by 0.45 s. Correct local execution or its validated timing budget before retest; do not waive the failed limit because temperature happened to stay low. Acquire the real OEM transient envelope and test-boundary approval. The OCP qualification framework supports testing at declared conditions; 13.5 owns integrated acceptance, including the other required pump, plant, restart and sensor states.
Ongoing fluid analysis: keeping the loop in spec for life
Commissioning hands you a clean, balanced, leak-tight loop. Operations has to keep it that way for five-plus years, and the loop degrades whether you watch it or not. PG25 and similar glycol coolants oxidize and lose inhibitor; dissimilar metals in cold plates, manifolds, and the CDU set up galvanic corrosion; biofilm grows wherever the biocide thins; and particulate sheds from every wetted surface to foul the microchannels that close your thermal budget. A fouled cold plate raises junction-to-coolant resistance — the largest term in the thermal ladder — and quietly eats the margin you commissioned. The countermeasure is a fluid-analysis program: scheduled sampling for pH, inhibitor concentration, conductivity, dissolved metals (the corrosion signature), and biological activity, with side-stream filtration (sub-micron on hyperscale CDUs) and biocide dosing tuned to the results. The cadence is a real cost and a real decision — use the supplier-approved interval and adverse-trend triggers established in 5.7 — and it is cheaper than the alternative, which is discovering corrosion by the leak it eventually causes.
This is the last of Part 5's mechanical content. The coolant chemistry and material-compatibility envelope is set in Chapter 5.4 (coolant selection) and Chapter 5.6 (CDU filtration and chemistry); the facility-water treatment and biocide/Legionella program for the heat-rejection side is in Chapter 5.7 and Chapter 5.8; the pressure-system mechanical engineering — pipe code, water-hammer, NDE, and hydrostatic acceptance that underwrites the L3 pressure test — is in Chapter 5.13; and the flush/fill/pressure acceptance procedures referenced here are detailed in the construction-execution path at Chapter 13.5.
Choose the local response and capacity arrangement that meets the named fault contract without a supervisory dependency. A missed heat-removal deadline requires a changed protection sequence or reduced operating load; a steady spare-capacity calculation alone does not preserve useful output through the transition.
Cite this chapter
Fehn, J. (2026). Thermal Design, Reliability, Leak Detection & Commissioning (Chapter 5.11). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-11-thermal-design-reliability-leak-detection-and-commissioning (accessed 2026-09-29).
@misc{aidc-5-11,
author = {Fehn, Jacob},
title = {Thermal Design, Reliability, Leak Detection & Commissioning (Chapter 5.11)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-11-thermal-design-reliability-leak-detection-and-commissioning},
note = {Accessed 2026-09-29}
}