Chapter 13.5
In this chapter · 7 sections
Cooling Acceptance: Air, Liquid-to-Chip & CDU Commissioning
Commission the cooling protection chain in layers: surrogate load and injection/HIL prove controls and fail-safe actions, liquid-cooled load banks on the manifolds prove the liquid path, and staged normal workloads close operating evidence without making live GPUs the fault target.
What you'll decide here
- Which of the three evidence layers — dry and surrogate-load tests, liquid-cooled load banks on the rack manifolds, the first staged workload — closes each acceptance item, and which protective-function tests stay on injection and an isolated test scope rather than live GPUs.
- The approved fluid-cleanliness criterion before first coolant touches a cold plate — conductivity, particles, compatibility, limiting-branch samples, and the flush cycles and resampling you budget — because under-flushing fouls microchannels you cannot clean in place and turns saved hours into a live-rack replacement campaign.
- Whether you witness the delivered CDU/control factory acceptance test (FWT) or accept on the datasheet alone, then re-prove installed duty — the fork that decides whether a pump-curve or control-loop defect surfaces in a factory bay or in your hall while the GPUs wait.
- How the hydraulically limiting branch meets minimum flow while every other branch draws its coincident duty, using rack/manifold fixtures before full population and confirming the production operating point afterward.
- How fast the leak-detection and loss-of-flow interlocks throttle or park the GPUs, proven by injected trips before silicon arrives — a leak is never staged on a live rack to prove the logic.
Every other acceptance domain in Part 13 can be exercised to its design point with surrogate load. Electrical acceptance drives the switchgear and UPS with load banks (Chapter 13.3); generators and microgrids are paralleled and islanded against resistive and reactive banks (Chapter 13.4); integrated systems testing uses approved loss-of-source simulation, secondary injection, HIL, or a controlled isolated-boundary transfer before any exceptional live high-energy test (Chapter 13.6). Cooling acceptance is the exception that defines the whole part. An air-rejecting facility load bank is a resistor stack with a fan: it converts megawatts into hot air and rejects that air into the room. It does not, and cannot, push heat through a cold plate into the secondary liquid loop. Liquid-cooled load banks can drive the CDU secondary loop and rack/manifold branches before production GPUs are racked; a staged workload then closes die/TIM/contact and workload-specific evidence.
That limit organizes the chapter: airside acceptance first, then the secondary-loop work that can be done dry or with surrogate heat — fill and vent the isolated test scope, hydrostatic acceptance, drain and dry, flush to fluid quality, coolant charge and purge — then CDU acceptance and the worst-case-branch problem, and finally the leak-integrity, failover, and burn-in interlocks that close the gate. Each fork carries a downstream cost that comes due when real silicon arrives. Mechanical commissioning and GPU burn-in (Chapter 13.8) are not adjacent phases with a clean baton-pass; they overlap, by physics, because a staged workload closes the hardware- and workload-specific evidence that liquid-loop emulation leaves open.
The acceptance map: what can be proven, and with what load
Cooling acceptance spans two physically distinct systems joined at the CDU. The facility water system (FWS) — chillers or dry coolers, towers, the primary loop, pumps, and the airside plant — is conventional mechanical Cx, and most of it can be driven to design with surrogate load: load banks dump heat into the room for the air handlers to reject, and the primary loop can be exercised by the CDU's own heat exchanger or by temporary process loads. The technology cooling system (TCS) — the secondary loop the CDU isolates from facility water, the in-rack manifolds, the quick-disconnects, and the cold plates themselves — is where the load-realism limit bites. You can fill it, pressure-test it, flush it, and run the pumps; rack/manifold-connected liquid-cooled load banks then impose controlled steady and transient heat while exercising secondary-loop flow, ΔT, control response, and branch hydraulics without production silicon. The CDU/TCS separation and loop architecture are engineered in Chapter 5.6; here we accept what was built there.
| Acceptance item | System | Provable pre-GPU? | Surrogate used | Evidence remaining after the pre-GPU test |
|---|---|---|---|---|
| Airside / room cooling capacity | FWS (air) | Yes — fully | Load banks reject to air | Production airflow, placement and coincident liquid/air duty still need confirmation |
| Primary loop, heat rejection, free-cooling changeover | FWS (liquid) | Yes — to design heat | CDU HX or process load | Real annualized climate sequencing over seasons |
| Flushing, fluid quality, fill/purge | TCS | Yes — must precede GPUs | Named OEM/material/warranty-approved flush fluid and procedure, then the approved working fluid | Long-term chemistry drift, biofouling onset |
| Hydrostatic / pressure-integrity test | Both | Yes — after fill-and-vent of the isolated pipe scope | Compatible test fluid + hydrostatic pump | Hot-state cycling, service-cycled QDs and transient surge need separate evidence |
| CDU flow, head, pump redundancy, loop transient stability | CDU | Yes — at rated flow | Pump-only circulation plus liquid-loop emulator | Hardware- and workload-specific operating evidence |
| Worst-case-branch flow at full load | TCS | Partially | Rack/manifold-connected liquid-cooled load banks | Production-rack and workload-specific confirmation |
| Production-silicon thermal response and workload interlocks | CDU + TCS | No | — | Measured die temperatures, throttling, and interlock timing under the staged workload |
| Leak-detection + cooling-failover interlock to throttle | TCS + IT | Partially | Manual trip injection | GPU throttle/park actually fires before Tj runaway |
Airside acceptance: the part that behaves
Even a fully liquid-cooled hall has a residual air load — 17 kW per HPE GB200 NVL72 rack stays on air (NICs, DIMMs, PSUs, optics, switch trays), and storage, networking, and any modest-density inference rows may be entirely air-cooled (Chapter 5.2). Lenovo Press LP2357 puts GB300 NVL72 on a 90/10 liquid/air heat split, roughly 13.5 kW on air at its 135 kW rack TDP and about 15.5 kW at the 155 kW peak, with the NVLink switch trays moved fully to liquid. Airside acceptance is the conventional, well-understood half of this chapter, and it is genuinely provable pre-GPU because air is the load bank's native sink. The work: verify CRAH/RDHx/in-row capacity at design heat with load banks placed to mimic the rack thermal map; commission containment (hot/cold-aisle or rear-door) for leakage and bypass; tune supply-air setpoints against the ASHRAE A1–A4 envelope; and prove airflow balance so no rack starves. For hybrid halls running DLC plus rear-door exchangers (Chapter 5.3), the RDHx water side is part of this acceptance and its condensation/dew-point margin is set here.
The decision that matters in airside acceptance is how much residual-air capacity you commission relative to the liquid fraction. Over-commission and you have paid for air-handling you will idle as the hall liquid-cools more of the load through the density ramp; under-commission and the room plant comes up short when the installed rack mix raises residual-air duty beyond the accepted envelope; cold-plate and flow faults must instead prove alarm, isolation, power capping or throttling, orderly shutdown, and trip because the liquid-cooled silicon has no alternate air path. Commission against the OEM-declared residual-air duty for the racks in scope, and preserve physical and electrical provisions for later airside expansion; total rack power does not define the future air load.
Secondary-loop flushing and fluid quality: the gate before first coolant
Before a drop of working coolant touches a cold plate, the secondary loop must be flushed and qualified, and this is the most under-budgeted step in cooling acceptance. The cold-plate microchannels that make DLC work — sub-millimeter passages that drive the convective coefficient — are precisely what particulate and biological fouling block, and once a cold plate is fouled you cannot clean it in place; you replace it, in a live rack, with the loop drained. The flush is the gate that protects the most expensive and least serviceable surface in the building. The acceptance criterion comes from the coolant supplier's handling instruction for the named fluid, and it is numerical: Arteco's March 2026 Freecor EV Micro 10 instruction, for example, recommends rinsing with that product or ultrapure water below 5 µS/cm, followed by a full drain before charging the working fluid (the flush cadence itself is Chapter 5.7). Hold the loop to the named fluid's own limits — conductivity, particulate class, residual glycol — sampled at the furthest branch rather than at the CDU, and record the wetted-material matrix, water source, fluid batch and concentration with the release signature; a different fluid carries a different number. → Chapter 5.4.
Hydrostatic and pressure-integrity acceptance
Pressure-integrity can be tested before GPUs because it tests the pressure boundary; a cold static hold does not accept hot-state cycling, QD service cycles or surge. The charged-piping code basis is ASME B31.x in North America or the EU PED / EN 13480 fork in Europe (Chapter 5.13), and hydrostatic acceptance begins by filling and venting the isolated pipe scope with a compatible test fluid. The governing code sets the hold pressure: B31.3 and B31.9 have different scope and testing rules; use the adopted code's pressure, temperature correction, examination and hold requirements for the exact pipe scope; components not rated for the pipe-scope test are removed or isolated and proven under their own OEM-qualified procedures, never a blanket multiplier. Use the required hold and examination sequence with stabilized fluid temperature; pressure decay needs temperature correction before it can identify a joint or gasket that seated only under initial pressure. The acceptance sequence is strict and ordered: fill and vent the isolated test scope with test fluid, pressure-test, drain and dry where required, flush to the cleanliness criterion, then charge the working coolant and purge. Reorder it and you either flush a loop you have not proven leak-tight or charge working fluid into a loop you have not cleaned.
The decision embedded here concerns the quick-disconnects. A GB200-class rack carries on the order of 150–200 dripless quick-disconnects, and every one is a potential leak path, but the isolated pipe-scope hold does not prove the QD population at a blanket multiplier. Remove or isolate any QDs not rated for the pipe-scope test and prove them under their own OEM-qualified procedures, then cycle a representative sample under the OEM's qualified procedure to catch couplings that seal statically but weep after a service cycle. The sample adds acceptance time; skipping it buys leak risk during the first board-swap. Given that serviceability is the whole point of dripless QDs, cycling a representative sample is the defensible call.
CDU acceptance: factory witness, flow verification, and the worst-case branch
The CDU is the seam of the entire cooling system — it isolates the technology loop from facility water, sets secondary flow and temperature, and carries the controls that must respond to load. It is also, as Uptime Intelligence has flagged, the component most likely to complicate commissioning, because many CDU vendors arrived from outside the data-center world and some had never integrated a unit into a complex fluid network before. That makes the factory witness test (FWT) decision consequential: witness the CDU's flow, head, pump-redundancy failover, and control response in the vendor's bay, or accept it on a datasheet and discover a defect in your live hall. The cost asymmetry is stark — a pump-curve or PID defect found at the factory is a vendor rework; found in the field it is a hall-level schedule hit with the cluster waiting. For any first-of-a-kind CDU model or vendor, FWT is the rational default.
On site, CDU acceptance proves rated flow and head, pump N+1 failover (kill the lead pump, confirm the lag pump holds flow without a thermal excursion), filtration and dew-point control, and the leak-detection integration. Flow verification follows the approved staged profile — for example 25% → 50% → 75% → 100% as an illustrative sequence, not a portable ramp rule — with temperature differential, flow rate and pressure drop logged across the CDU, piping and rack manifolds at coincident heat duty. But that staging runs against surrogate or balancing-valve load, which brings us to the hardest problem in the chapter.
The worst-case branch. A liquid loop balances flow across many parallel branches; the branch with the least available differential pressure relative to its required flow is the hydraulic candidate for starvation at full load; the hottest branch can be elsewhere, so instrument both. Acceptance practice instruments that worst-case branch and verifies it makes its minimum flow when the whole system is loaded. Connect liquid-cooled load banks at every rack/manifold branch in the acceptance scope to impose controlled simultaneous steady and transient heat; use balancing valves to reproduce the branch with the least differential-pressure margin at required flow, and instrument the hottest branch separately. This proves branch flow, balance, and control response before production silicon. A staged workload then confirms production-rack and workload-specific operation.
| Acceptance item | Liquid-bank / injection result | What the first staged workload re-reads | Consequence of skipping the re-read |
|---|---|---|---|
| CDU rated flow & head | Proven at 100% flow | — | Fluid viscosity, temperature, branch resistance and coincident duty change the operating point |
| Pump N+1 failover | Proven (kill lead pump) under liquid-bank heat | Repeat at the production operating point: flow held is not junction margin held | Failover may hold flow but not Tj margin under real heat |
| Worst-case-branch flow | Forced to minimum flow with every branch loaded by liquid-cooled banks | Re-read at full population under the first workload | A branch can starve under coincident demand, pump loss or a restrictive valve/filter state |
| Control-loop / setpoint stability | Proven against programmed heat steps on the liquid bank | Re-check at the production operating point; retune if it hunts | Hunting, oscillation, or dew-point excursion in production |
| Leak-detect to GPU-throttle interlock | Sensor trip injection: retain BMC throttle/park event and measured actual power/heat reduction against the protective deadline | Timing re-read with real racks in the loop — still by injection | Interlock fires too slow and Tj runs away on a real loss |
Scope & caveats
Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.
Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.
Scope & caveats
Guide water-property heat balance for the HPE 115 kW liquid load at a 10–7 K operating rise: density 1.00 kg/L and heat capacity 4.18 kJ/(kg·K). A derived design illustration, not an OEM flow requirement. The separate QCT 45 °C inlet and 65 °C return maxima do not prescribe this rise.
The operating flow must also satisfy the selected rack and CDU pressure/flow envelope.
Scope & caveats
The fuel-cell-coolant leaflet recommends rinsing with the product or ultrapure water below 5 µS/cm, then draining fully. This is a named-product instruction; use in a data-center loop requires the approved fluid and equipment compatibility basis.
Not a generic DLC flush or acceptance limit. Other fluids and equipment require their own approved compatibility, cleanliness, chemistry, and documentation criteria.
Scope & caveats
Reported forecast estimate, not a measured deployment census or a project cooling-selection rule. The cited PMR cold-plate category is broader than single-phase DTC.
Reported forecast estimate, not measured fleet share; PMR's published cold-plate category is broader than single-phase DTC and is not a project-selection rule.
Leak integrity, cooling failover, and the interlock with burn-in
Leak detection in a liquid-cooled hall is a real-time interlock that must throttle or park the GPUs before a coolant loss drives junction temperatures past their limit. Two architectural choices set the acceptance work. First, positive vs. negative-pressure operation: a negative-pressure (sub-atmospheric) secondary loop draws air in on a breach instead of pushing coolant out, turning a spray onto live electronics into an air ingress — a fundamentally safer failure mode that some designs adopt specifically to de-risk leaks. Acceptance must confirm the pressure regime behaves as designed under a fault. Second, the detection-to-action chain: rope/spot leak sensors, flow and pressure anomaly detection, and the logic that converts a detection into a GPU power-cap, throttle, or park. Acceptance injects faults — manual trips, simulated sensor alarms, a forced CDU pump loss — and confirms the action fires fast enough.
The protection chain is proven by injection: trip the sensor, spoof the alarm, kill the lead pump under liquid-bank heat, and record the BMC command and time the actual load reduction. The number to hold is the time to actual heat reduction or shutdown against the local limit in Chapter 5.11, first measured on the emulator and then verified with the installed control path; timing a BMC command alone leaves actuator response unproven. A leak is not staged on energized silicon to prove the logic; if an owner ever authorizes a bounded live fault test it is a documented exception with its own hazard analysis, not an acceptance step.
C-01 separate trace records. P starts from 4×80 =320 kW liquid and 40 kW air. At t =0 stop the lead pump; standby duty at 1.00 s and the stated P extrema cover every branch through t =9 s. Lower flow =112×0.980 =109.76 L/min (about 110); upper temperature =36.4 °C; pressure =0.09–0.51 MPa(g); standby upper time =1.05 s ≤1.20 s. P-R restores the original pump assignment/balance at 9 s (upper 9.05 ≤10), then records the declared baseline through t =39 s. Restore before the next injection.
D has no failed pump or leak input: request 60→80 kW on all four branches over t =0–1 s, so liquid heat rises 240→320 kW while air remains 40 kW. Lower flow =116×0.980 =113.68 L/min (about 114); upper temperature =36.2 °C; pressure =0.11–0.49 MPa(g). D-R returns to the declared 80 kW/branch baseline at t =9 s, within the 10 s recovery gate, and holds through t =39 s. This trace tests the load increase; it cannot be borrowed from P.
L starts again at 320 kW liquid plus 40 kW air. At t =0 inject branch-1 leak and request +20 kW there. The simulated control record refuses the increase, with actual branch heat never above 80 kW; detection at 0.10 s is distinct from actual heat off at 1.80 s and isolation at 3.00 s. Heat-off upper bound =1.85 s ≤2.00 s; surviving three branches retain P's bounds at 80 kW each. The isolated, heat-off branch has zero flow by design, not a passing loaded-flow result. L-R keeps it off until the deliberate reset at 10 s after simulated repair, leak-clear and restored baseline flow; then a 4 s ramp restores its 80 kW. At 14 s all branches regain baseline, within 10 s of reset, and hold through 44 s. No automatic restart is assumed.
Result and flip. P/P-R, D/D-R and L/L-R pass only their stated simulation gates. Their baseline lower flow is 117.6 L/min, upper temperature 35.2 °C and pressure 0.19–0.41 MPa(g). Whole-CDU release remains HOLD for the absent air-duty, sensor-failure and supervisory-loss records, actual fluid/pump curves, high-rate surge, calibrated power/thermal traces and witnessed reset. Mechanical owns branches/valves; platform/electrical owns actual power and command timing; operations and CxA own restoration/release. Abort a missed limit or lost critical channel using the independent heat-off path. Indicated heat-off crossover is 1.95 s; 2.40+0.05 =2.45 s fails by 0.45 s even if temperature stays low. Repair the actuator before increasing load. OCP's CDU methodology (August 2024 cover) supports declared-condition tests; 5.11 owns local limits. Saving the witness slot moves a pump/control defect into the first hot hall.
Deep dive: what liquid-bank and first-workload evidence each proves
A resistive facility load bank reproduces heat magnitude but not the die-to-cold-plate path; a manifold-connected liquid-cooled load bank reproduces the loop path but not the die/TIM/contact stack or the millisecond flux of a real die. The first staged workload (Chapter 13.9) closes what is left: supported supply and return conditions at the declared operating point, stable flow through the limiting branch, and no thermally attributable throttling under the real collective. Any acceptance item still open after that run is named, owned, bounded and signed by the owner, OEM and commissioning authority — a deferred item, not a live fault test.
Deep dive: the acceptance sequence, ordered (and why order is load-bearing)
Cooling acceptance is one of the few domains where getting the order wrong silently invalidates downstream tests. The defensible sequence:
- 1. Pressure-integrity / hydrostatic — fill and vent the isolated pipe test scope with compatible test fluid; hold at the governing-code pressure under the adopted B31.3, B31.9 or other governing pipe code; then drain and dry where required. Remove or isolate components not rated for the pipe-scope test and prove them under their own OEM-qualified procedures; cycle a representative QD sample under the OEM's qualified procedure.
- 2. Flush to the named acceptance specification — the coolant supplier's handling instruction sets the number (below 5 µS/cm for ultrapure rinse water in Arteco's March 2026 Freecor EV Micro 10 instruction); sample at the furthest branch and hold until it stabilizes, not until the schedule says so.
- 3. Charge working fluid & purge — the specified OEM-approved coolant and concentration, air-purge the loop, sample and label the fluid; trapped air destroys pump performance and the convective coefficient.
- 4. CDU acceptance at rated flow — flow, head, N+1 pump failover, filtration, dew-point control; FWT first for any new model.
- 5. Static worst-case-branch verification — use valves and dummy heat to reproduce least differential-pressure margin at required flow; confirm minimum flow and separately check the hottest branch.
- 6. Bounded failover & interlock evidence — kill the lead pump and inject the leak and flow-loss trips with the liquid banks running; time actual load reduction against the product's protective-action deadline.
- 7. Product-representative staged evidence — read the operating point, the limiting branch and the throttle behavior under a real synchronized step; re-read the interlock timing by injection with the real racks in the loop.
The first six are mechanical Cx; the seventh is the burn-in overlap. Skip the ordering — flush before pressure-test, charge before flush, accept the CDU before the loop is clean — and each violation contaminates the step it precedes. The order is the dependency graph of the physics.
Anti-patterns
The recurring cooling-acceptance failures all share a root cause: treating an air-rejecting load bank as proof of the liquid loop, or treating mechanical Cx and burn-in as cleanly separable. Four are worth naming:
- Signing off cooling on facility load banks alone. Facility load banks prove FWS capacity, not the cold-plate path. Put liquid-cooled banks on the manifolds and read the first workload before signing the liquid side.
- Under-flushing to recover schedule. Calling the loop clean before conductivity and particulate truly stabilize. The debt is paid as a cold-plate replacement campaign in a live hall, weeks later, after unexplained throttling.
- Accepting a first-of-kind CDU on a datasheet. Skipping factory witness on a new CDU model or vendor. A pump-curve or control defect that would have been a factory rework becomes a hall-level schedule hit with the cluster idle.
- Using live silicon as the protective-function test fixture. The interlocks are proven by injected trips with the liquid banks running; the installed-rack check measures actual power reduction as well as command timing. A leak staged on a live rack proves nothing the injection did not, at the price of the rack.
Cite this chapter
Fehn, J. (2026). Cooling Acceptance: Air, Liquid-to-Chip & CDU Commissioning (Chapter 13.5). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-5-cooling-acceptance-air-liquid-to-chip-and-cdu-commissioning (accessed 2026-09-29).
@misc{aidc-13-5,
author = {Fehn, Jacob},
title = {Cooling Acceptance: Air, Liquid-to-Chip & CDU Commissioning (Chapter 13.5)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-5-cooling-acceptance-air-liquid-to-chip-and-cdu-commissioning},
note = {Accessed 2026-09-29}
}