The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 5.4

In this chapter · 6 sections
Term help

Direct-to-Chip Liquid Cooling (DLC) — Architecture, Fluid & Thermal Design

DLC is the design basis where the named equipment requires liquid capture at source; single- versus two-phase, coupling architecture, coolant, flow, delta-T, pressure drop, residual-air duty, and service/redundancy choices then set the lifetime operating and capex consequences.

POWER-BOUNDDENSITY-RAMPGOODPUT

What you'll decide here

  1. Which named single- or two-phase cold-plate system has OEM, fluid, warranty, service, safety and facility-envelope evidence for the project; reported market share does not select it.
  2. How you plumb the rack: blind-mate/floating-tray manifold couplings vs hand-mated flexible-hose quick disconnects — the access, tolerance and leak-risk fork that governs every node swap over the service life; either can self-seal, so qualify pressure loss, spill, air inclusion and the replaceable boundary for the selected tray.
  3. The per-chip and per-rack thermal budget—liquid-captured heat, approved-fluid properties, selected delta-T, and pressure drop across cold plate, manifold and QDs—with flow derived from the named operating point rather than a universal ratio.
  4. Coolant chemistry—the named OEM-approved base fluid, concentration, inhibitor package, water quality, wetted-material matrix, freeze/biofouling plan and warranty consequences across the secondary loop.
  5. What deliberately stays on air in the selected configuration — NICs, DIMMs, PSUs, optics or VRMs — and how its residual air duty is captured during service and faults, because a hall that forgets the air tail strands the very racks it cooled.
Illustrative — stated assumptions. One terminal arrowhead on each supply and return path shows fluid direction; dashed arrows show heat transfer. The residual-air label sits directly on its dashed path. TCS means technology cooling system and FWS means facility water system. Isolation valves and quick disconnects support the selected service procedure; neither symbol proves safe disconnection under pressure or load. Close duty, inlet, fluid and hydraulic conditions in the required degraded state.

By 2026 the argument about whether AI racks need liquid is over. GB300 NVL72 shipped in 2025 and deploys through 2026 alongside GB200; Lenovo specifies 135 kW rack TDP and up to 155 kW peak. The best-documented published heat split is still the cited HPE GB200 NVL72 record, which assigns roughly 115 kW to liquid and 17 kW to residual air (Chapter 5.2); this chapter derives from that record and scales to the rack you name. That product-specific heat split makes direct-to-chip liquid cooling (DLC) plus residual room-air removal the supported basis; Chapter 5.1 explains the wider selection envelope. Rear-door heat exchangers and air-assisted liquid (Chapter 5.3) extend a brownfield envelope only where the named product rating, airflow or room-rejection path, water conditions, service/redundancy case, and refresh tail close. For a greenfield whose selected or roadmap rack profile requires liquid capture at source, DLC and its residual-air path belong in the design basis before steel is cut. For VR200 NVL72, size irreversible infrastructure to NVIDIA's 330 kW cabinet TDP facility design basis and run energy and TCO models on Pegatron's 188 kW Max Q / 228 kW Max P operating profiles; NVIDIA reported Vera Rubin in full production in August 2026.

Once DLC is assumed, the decisions that remain are concrete. Choose two-phase with a refrigerant that meets the applicable PFAS definition and you inherit that supply-chain and liability problem alongside the pressure, vapor and OEM-support obligations. Choose hand-mated flexible hoses over blind-mate manifolds and you trade factory-aligned tray geometry for field access and routing freedom — and a different leak-risk profile; either can self-seal, so qualify spill, cycle life and air inclusion for the actual service procedure. Budget the cold-plate delta-T too tight and you over-spec the CDU and the pumps; too loose and you throttle the GPUs. Each of these forks is cheaper to see before it is poured into a slab.

Cold-plate architectures: single-phase vs two-phase

A direct-to-chip cold plate is a sealed metal block — typically copper, sometimes copper-on-aluminum — pressed onto the die package through a thermal interface material (TIM), with coolant forced through internal microchannels or skived fins directly over the hot silicon. The heat path is short and the thermal resistance low: a few hundredths of a °C per watt from junction to coolant, which is what makes 1.0–2.3 kW per GPU package tractable. The first fork is whether the coolant changes phase inside the plate.

Single-phase cold plates keep the coolant liquid throughout. A water/glycol mix enters, warms by a bounded delta-T (commonly 7–12 °C), and leaves still liquid. Heat removal scales with mass flow times specific heat times delta-T — pump harder or run a wider delta-T to carry more watts. It is mechanically simple, the fluids are benign (water-based, non-PFAS), the pressure regime is well understood, and it maps cleanly onto the CDU-and-secondary-loop architecture of Chapter 5.6. The penalty is that water has finite heat capacity, so very high heat fluxes demand high flow and therefore pump energy and pressure drop.

Two-phase cold plates exploit the latent heat of vaporization: a low-boiling-point dielectric enters as liquid, boils inside the plate, and leaves as a vapor-liquid mixture. Because latent heat dwarfs sensible heat, two-phase moves enormous heat flux at low flow and a nearly isothermal plate surface — a thermally attractive candidate for the 1.5–2.3 kW package design problem when its component temperature, flow stability and rejection conditions beat the supported single-phase alternative. The catch is the working fluid. The engineered dielectrics that boil at convenient temperatures have so far included fluorochemicals in the PFAS family under the broad EU proposal’s definition, and that is a regulatory and liability exposure rather than a footnote; the 2026 HFO alternatives are fluorinated too, so qualify the named fluid’s composition, warranty and supply instead of trusting the phase label.

Single-phase vs two-phase direct-to-chip — the cold-plate fork
AxisSingle-phase DLCTwo-phase DLC
Heat-transfer modeSensible heat; coolant stays liquidLatent heat; coolant boils in the plate
Working fluidNamed OEM-approved water or water/glycol fluid; concentration and inhibitor package are project-specificEngineered dielectric — fluorochemicals, PFAS under the broad EU definition
Flow demandDerived from liquid heat, approved-fluid properties and selected ΔT; no universal L/min/kWMuch lower; latent heat does the work
Plate surface tempRises by the selected operating ΔT inside the named product envelopeNear-isothermal at the boiling point
Pressure regimeUse the named OEM pressure-flow curve; no portable per-plate valueTwo-phase flow instability risk; harder to control
2026 status~55% in one 2026 market estimate; not a project-selection ruleNamed pilots and vendor evidence; production adoption remains gated by OEM/warranty, fluid and service evidence
Primary riskPump energy/flow at very high fluxFluid supply chain, regulation, liability
Design bands are 2026-current practitioner figures (DCD/Schneider, OCP, Dober, SemiAnalysis). Two-phase figures are pilot/early-deployment, not mature production.

In-rack plumbing: manifolds, blind-mate, and quick disconnects

Getting coolant from the rack inlet to 72 cold plates and back, while letting a technician swap a failed tray in minutes without draining the rack, is the second major fork in DLC design. Every NVL72-class rack carries a vertical pair of manifolds (a supply and a return rail, the in-rack analogue of a busbar) running the rack height. Each compute tray taps the rails through couplings. The decision is what kind of coupling, and it trades serviceability against reliability and leak risk.

Blind-mate / floating-tray couplings are integrated into the tray and the manifold so that sliding the tray home automatically engages the fluid connection — no hose to route, no fitting to hand-torque. The 'floating' geometry absorbs the mechanical tolerance stack so the connection self-aligns. This is the factory-integration path: couplings are validated at L10/L11 integration (Chapter 5.13 on the mechanical side; rack integration in Part 7), and field service becomes a slide-out/slide-in operation. The cost is rigidity — the rack and tray geometry are co-designed and far less forgiving of field improvisation.

Hand-mated flexible-hose quick disconnects put a short hose with a dry-break coupling between the tray and the manifold. The technician physically connects two halves; the dry-break valve seals both sides on disconnect so the spill is bounded to the qualified small release rather than an open stream, with both spill and air inclusion checked at the specified pressure and cycle count. This is more serviceable in the messy reality of a live hall and tolerant of tolerance stack-up, but every manual connection is a potential leak point and a human-error surface, and the hoses add pressure drop and clutter. OCP has standardized UQD form factors precisely to make these field-mateable and second-source-able.

Require the supplier to identify the exact connector, manifold and rack-interface revisions and the tested mating pair. The OCP Cold Plate index separates accepted contributions from work in development; the Open Rack index identifies the rack documents. An OCP interface name does not establish cross-vendor interchangeability, seal compatibility or warranty. Freeze those interfaces in the purchase schedule and rehearse the actual tray replacement before approving the service boundary.

Per-chip and per-rack thermal design

The thermal budget is governed by one conservation equation: the heat a loop carries equals mass flow times specific heat times the coolant temperature rise (Q = ṁ · cp · ΔT). Everything in DLC design is a negotiation among the three terms on the right — and against a pressure-drop ceiling that the pumps and CDU must overcome.

Flow from the heat balance. Use Q = ṁ·cp·ΔT on the liquid-captured heat. For water-like properties, 1 kW at a 10 K rise requires about 1.43 L/min; recalculate with the approved fluid at its operating temperature. For the HPE GB200 NVL72's ~115 kW liquid share, water-property illustrations are about 82 L/min at 20 K, 165 L/min at 10 K, and 236 L/min at 7 K. A Dober PG25 planning illustration of roughly 1.25–2.0 L/min/kW at 7.5–12 °C is supplier guidance, not an OCP or OEM acceptance requirement. Size the manifold and CDU from the named vendor schedule or declared-fluid calculation plus the pressure and control margin. Run a wider delta-T and you cut the flow (and pump energy) for the same watts — but you raise the return-water temperature the heat-rejection plant must handle and you push the warmest cold plates closer to the throttle line.

Wide vs. tight delta-T. A wider coolant delta-T is a gift to the facility: it means less flow, smaller pipes, lower pump energy, and warmer return water that free-cooling and heat-reuse plants love (Chapter 5.7, Chapter 5.9). For the cited QCT GB200 NVL72 reference, 45 °C maximum liquid inlet and 65 °C maximum liquid return are separate limits, not a prescribed 20 K operating rise; ASHRAE W45 describes FWS supply capability. The junction budget still binds: the delta-T you can run is bounded by the supply temperature plus the cold-plate's thermal resistance, so spend the budget on a wide delta-T at a warm inlet and the last plate in a series path may sit too warm. Design teams therefore favor parallel manifold paths so every cold plate sees near-inlet coolant. Series routing cuts the total flow the manifold must carry and simplifies the run, but plate pressure drops add along the string and every downstream chip starts warmer — so choose it only where both the required head and the last plate's junction temperature still close at the named per-plate flow.

Pressure-drop budget. Pump head is divided among the selected cold-plate pressure-flow curve (including declared fittings and quick disconnects), the in-rack manifold and headers, and the secondary loop back to the CDU. Rack-level pressure drop runs through the headers, hoses, QDs and cold plates on the limiting parallel path; use the assembly pressure-flow curve and name what it includes, or counting the same loss twice will oversize the pump. Every fitting, every hose, every reduction in channel size buys lower thermal resistance at the price of more pressure drop — and pressure drop is pump energy, which shows up in PUE. The cold-plate designer's lever — finer microchannels for lower thermal resistance — is exactly the lever that raises pressure drop, so the per-chip design is an explicit optimization of thermal resistance against pressure drop.

Choose series or parallel from the last component’s margin

Parallel routing needs 4.00 L/min at 20.0 kPa plate-path loss. Both junctions are 35.0 + 1,000 × 0.0400 = 75.0 °C, retaining 10.0 K. Series needs 2.00 L/min at 40.0 kPa; the first plate raises water temperature by 1.00 × 60/(2.00 × 4.18) = 7.18 K, so the second junction is 82.2 °C and retains only 2.82 K. Choose parallel to preserve the required reserve. Ideal plate-only hydraulic power is 1.33 W in either arrangement; pump efficiency and external losses decide electrical input.

The series choice reverses at supply ≤85.0 − 5.0 − 40.0 − 7.18 = 32.8 °C, calculated without intermediate rounding. Colder supply spends rejection opportunity to avoid extra header flow. Alternatively, qualify a different component/plate duty. OCP’s cold-plate method keeps the coolant-inlet reference explicit; use 5.1 for the thermal method and 5.13 for the complete hydraulic selection.

~55%forecast
reported 2026 forecast for single-phase cold-plate/direct-to-chip share; PMR's published cold-plate category is broader, and market share does not select a project architecture
Scope & caveats

Reported forecast estimate, not a measured deployment census or a project cooling-selection rule. The cited PMR cold-plate category is broader than single-phase DTC.

Reported forecast estimate, not measured fleet share; PMR's published cold-plate category is broader than single-phase DTC and is not a project-selection rule.

~1.43 L/min/kWderived
guide water heat-balance at 10 K: ~1.43 L/min per kW of liquid-captured heat; recalculate for the approved fluid
2025Guide heat-balance derivation using water properties at the declared design pointregister ↗
Scope & caveats

A guide physics result, not an OCP or OEM flow specification. Apply it only to liquid-captured heat and recalculate with the approved fluid properties, selected ΔT, pressure budget, and product operating envelope.

45 °C maximum liquid inlet; 65 °C maximum liquid return (separate limits)
QCT GB200 NVL72 QoolRack reference maxima: 45 °C liquid inlet and 65 °C liquid return; separate limits, not a selected operating pair
Scope & caveats

Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.

Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.

~165–236 L/minderived
HPE liquid-load water-property example at 7–10 K rise; Chapter 5.1 owns the heat-balance method, and the selected rack/CDU envelope sets allowable flow
Scope & caveats

Guide water-property heat balance for the HPE 115 kW liquid load at a 10–7 K operating rise: density 1.00 kg/L and heat capacity 4.18 kJ/(kg·K). A derived design illustration, not an OEM flow requirement. The separate QCT 45 °C inlet and 65 °C return maxima do not prescribe this rise.

The operating flow must also satisfy the selected rack and CDU pressure/flow envelope.

~115 / ~17 kW
HPE GB200 NVL72 heat split — 115 kW liquid / 17 kW air at 132 kW nominal
Scope & caveats

HPE GB200-specific. Keep separate from the GB300 NVL72 split: Lenovo Press LP2357 puts GB300 on ~90% liquid / ~10% air — roughly 13.5 kW on air at 135 kW rack TDP and ~15.5 kW at the ~155 kW peak, with the NVLink switch trays moved fully to liquid. Size a residual-air path from the rack you actually name.

$300–500/kW
Direct-to-chip system capex — a historical secondary estimate for its stated scope; obtain a matched installed-project estimate before selecting cold plates
up to 3x
in-silicon microfluidic cooling vs cold plates (≤65% lower peak temp rise) — the forward pointer
Deep dive: coolant selection — why PG25, and the consequences of the choice

The secondary-loop coolant is a chemistry decision with mechanical consequences that ripple from the cold plate to the CDU. There is no universal single-phase coolant recipe. Select the rack-side fluid from the accelerator, cold-plate, CDU, quick-disconnect, seal, pipe, heat-exchanger, climate, water-quality, and warranty requirements. PG25 — a 25% propylene-glycol/water blend with an approved inhibitor package — is one project-specific option, not the default for every loop.

Heat transfer. Pure water has the best specific heat and lowest viscosity — thermodynamically you would run water if you could. Glycol degrades both: it raises viscosity (more pump energy, more pressure drop) and lowers specific heat (more flow for the same watts). Use only the concentration that the freeze case, material-compatibility matrix, biological-control plan, fluid supplier, and equipment warranties support.

Freeze and biofouling protection. The glycol you do add buys freeze protection for outdoor loop sections and dry coolers in cold climates, and propylene glycol (vs ethylene) is chosen for low toxicity — a leak near electronics and people is less hazardous. Biological control is non-negotiable where the approved loop program requires it: warm water can grow biofilm that fouls microchannels and spikes pressure drop. Use the fluid supplier’s analytical program; add a separate biocide only with its approval. Temperature rise and flow follow the named equipment profile, approved-fluid properties, heat load, pressure budget, and control range; a generic PG25 ratio is not an acceptance criterion.

Material compatibility. The loop is a mixed-metal system — copper cold plates, stainless or brass fittings, aluminum heat exchangers, EPDM/elastomer seals. Galvanic corrosion and incompatible elastomers are the failure modes that surface late: the wrong inhibitor package or an unmanaged pH lets dissolved copper plate out on aluminum and seals swell or embrittle. Untreated deionized water can be aggressive to some metals and supplies neither biocide nor freeze margin, so fluid chemistry must follow the complete OEM-approved wetted-material and water-quality specification. The coolant is therefore a managed fluid: filtration, periodic chemistry sampling, and inhibitor top-ups are an operational line item, not a fill-and-forget. → the canonical fluid envelope and batch decision are in Chapter 5.7; the CDU implements the specified filtration and monitoring duties.

What stays on air — and how it is handled

'Liquid-cooled' overstates what DLC does. On the GB200/GB300-class racks documented here it cools the high-flux components only — the GPUs, the CPUs/Grace dies, the NVLink/NVSwitch silicon, and increasingly the high-power VRMs — leaving a residual air load that a hall ignores at its peril. On a GB200 NVL72 rack, roughly 115 kW is removed by liquid and ~17 kW remains on air — about 13% of the rack. Scale the split with the rack you name: GB300-class racks run higher total power and a higher liquid duty, so an air tail sized off the GB200 record is sized for the previous generation — and the tail is not permanent, because NVIDIA's Vera Rubin generation is specified fully liquid-cooled with the rack fans removed and the NICs, optics and power boards on the loop (Chapter 5.2). That tail is everything not worth a cold plate: NICs and optical transceivers, DIMMs, power-supply units, lower-power voltage regulators, the BMC, and miscellaneous board components. Optics in particular are a growing concern — pluggable transceiver power is climbing, and the optics sit at the rack's air-cooled edge precisely where airflow is now sparse.

The fork here is how you capture the air tail. Three patterns dominate. In-rack air-to-liquid: a small rear-door or in-chassis air-to-liquid heat exchanger rejects the residual air load back into the same liquid loop, so the rack exhausts neutral air and can make a ‘zero-air-to-room’ rack possible when coil temperature and airflow close; the hall needs no separate mechanical air plant only if another sink or demonstrated heat reduction covers the required failure-state duty. Hybrid containment: the hall keeps a reduced CRAH/in-row air system sized only for the ~10–20% air tail, with hot/cold-aisle containment, which is simpler to retrofit but reintroduces an air plant and its PUE. Facility air: let the tail dump into the room and pay for the building air system to catch it — distribution, installed capacity and degraded operation must close at the actual rack count; compare that plant cost with rack coils. A hall that sizes liquid for 115 kW and forgets the 17 kW air tail will thermally throttle on the optics and DIMMs while the GPUs run cold — stranding the rack it just paid to liquid-cool. → containment strategy for hybrid halls in Chapter 5.3; Chapter 8.3 qualifies the populated switch and port-bank cooling envelope; 8.9 owns optical-channel qualification and 8.10 the packaging and replaceable-part boundary. Liquid cooling of a switch ASIC does not qualify its faceplate optics.

What gets liquid, what stays on air — and why
ComponentCooling pathRationale
GPU / accelerator packageLiquid (cold plate)High source heat flux; requires the selected component cooling interface
CPU / Grace dieLiquid (cold plate)High flux; co-located on the tray
NVSwitch / NVLink siliconLiquid (cold plate)Dense interconnect silicon, significant draw
High-power VRMsLiquid (increasingly)Power-delivery losses now warrant a plate
NICs / optical transceiversAirLower flux but rising; sit at the rack edge
DIMMs / memoryAir (often)Distributed, lower flux; hard to cold-plate
PSUs / BMC / miscAirFollow the named configuration; some products put these components on liquid
Representative NVL72-class split; exact allocation varies by vendor and generation. Air tail ~10–20% of rack power.

Forward pointer: in-silicon microfluidics

Cold plates have a hard physical limit: no matter how good the plate, heat must still conduct from the junction, through the package, across the TIM, and into the plate before the coolant ever sees it. That stacked thermal resistance — and especially the TIM — is what caps the flux a cold plate can handle. The next step removes the intermediary entirely: etch the coolant channels into the silicon itself.

In-chip (direct-to-silicon) microfluidics routes coolant through microscopic channels — each roughly a hair's width — cut directly into the die or the backside of the package, so liquid flows over the hotspots inside the chip rather than across an external plate. Microsoft's 2025 prototype, using AI-designed, leaf-vein-inspired channel networks, reported up to 3x better heat removal than state-of-the-art cold plates and up to 65% lower peak temperature rise. The rationale is pure thermal resistance: collapsing the conduction path from junction to coolant is the only way to keep pace with 3D-stacked dies and the projected 2–3 kW packages behind them, where a cold plate simply runs out of room. Treat it as a roadmap signal rather than a 2026 production technology. The full consolidated cooling roadmap, including immersion's role and the 600 kW–1 MW rack generation, lives in Chapter 16.2.

DLC sits in the middle of Part 5's cooling stack. The density wall that forces it is in Chapter 5.1; the air regime it replaces in Chapter 5.2; the RDHx/AALC bridge for brownfields in Chapter 5.3; immersion's parallel single-/two-phase story and the PFAS reckoning in Chapter 5.5. The secondary loop that feeds these cold plates — CDUs, fluid chemistry, dew-point margin — is Chapter 5.6; the facility water loop and warm-water strategy Chapter 5.7; heat rejection Chapter 5.8; heat reuse Chapter 5.9; retrofitting air halls to liquid Chapter 5.10; reliability, leak detection and commissioning Chapter 5.11; and the mechanical/pressure-system engineering of the piping Chapter 5.13. The archetype decision that made DLC mandatory is framed in Chapter 1.1; the in-silicon microfluidics roadmap in Chapter 16.2; and the optics thermal tail in Chapter 8.10.

Choose the plate topology, connection interface and service isolation together. Series plumbing saves flow only while the downstream component retains its margin; parallel plumbing spends distribution capacity to preserve a common inlet. Release procurement against the complete supported assembly and its qualified operating point.

Cite this chapter
Fehn, J. (2026). Direct-to-Chip Liquid Cooling (DLC) — Architecture, Fluid & Thermal Design (Chapter 5.4). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-4-direct-to-chip-liquid-cooling-dlc-the-2026-default (accessed 2026-09-29).
@misc{aidc-5-4,
  author       = {Fehn, Jacob},
  title        = {Direct-to-Chip Liquid Cooling (DLC) — Architecture, Fluid & Thermal Design (Chapter 5.4)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-4-direct-to-chip-liquid-cooling-dlc-the-2026-default},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit