The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 5.1

In this chapter · 6 sections
Term help

Thermal Fundamentals & the Density Wall

Cooling is selected from the named rack and facility envelope: rack kW is one input alongside liquid/residual heat split, component heat flux, airflow and inlet limits, TCS/FWS conditions, room rejection, climate, serviceability, redundancy, and the refresh tail.

POWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. Which named current and roadmap rack profiles the facility must support, including their heat split and flux, airflow/inlet envelope, TCS/FWS interfaces, residual-room duty, and service/redundancy case—because those requirements set the hard-to-retrofit structure, distribution, and heat-rejection provisions.
  2. Which air, RDHx/AALC, hybrid, or direct-liquid candidates close the complete equipment-and-facility envelope, and at which stated air/water temperatures, flows, pressures, fan state, and heat-capture target each product capacity is valid.
  3. The junction-to-coolant thermal-resistance budget you are designing against — the chip vendor fixes the temperature limit and heat load for the selected operating mode, leaving you the qualified coolant temperature and resistance stack to spend; name each resistance’s reference so the cold plate keeps the silicon legal without spending the same margin twice.
  4. The approach temperature and effectiveness you target at every heat exchanger in the chain, because each delta-T you spend narrows the free-cooling window and pushes the facility toward mechanical chilling.
  5. Whether the irreversible substrate (floor loading, facility water, pipe-rack and CDU space, electrical headroom) is sized for the density ramp, even where the reversible IT fit-out is matched to the current generation.
Cooling paths overlap. Qualify each against the named rack heat split and flux, airflow/inlet limits, water conditions, residual-room rejection, climate, service/redundancy case, and refresh tail before fixing the hall's distribution and structure.

Every accelerator is, thermodynamically, a space heater that happens to do arithmetic. A 1,200 W GPU converts essentially all of its electrical input into heat, and that heat must be removed continuously and within a few degrees of a fixed temperature limit or the silicon throttles, ages, or fails. This is the one constraint in the building that does not negotiate. You can oversubscribe a fabric, defer a redundancy tier, re-price a power contract — but you cannot argue with the second law. The heat leaves through the path you built for it, at the rate physics allows, or the machine slows down to match the path you actually have.

Part 5 builds on this thermal foundation. It starts with heat-flux first principles and the thermal-resistance stack-up from junction to coolant, then follows the 2020–2027 rack roadmap while testing conventional air, close-coupled air, RDHx/AALC, hybrid, and DLC against named equipment and facility envelopes. The thermal metrics—approach temperature, NTU/effectiveness, airflow and pressure, coolant delta-T, flow, and pressure drop—become the qualification vocabulary. Rack kW narrows the feasible set; heat split and flux, airflow/inlet limits, FWS/TCS or entering-water conditions, room rejection, climate, service/redundancy, and the refresh tail select the service.

Heat flux and the resistance stack-up

The governing quantity is not power but heat flux — power per unit area, W/cm². A 700 W H100 SXM module, with its full rated power hypothetically spread over roughly 8 cm², gives 85–90 W/cm² by division — a guide scaling calculation from NVIDIA’s H100 rating, not a measured die heat flux. The module’s power includes heat outside the die, and a Blackwell-class package past 1 kW still needs its own cooled area and spatial heat map before you call an average a hotspot. Flux is what the cooling solution actually fights, because heat removal is fundamentally limited by how much surface area you can couple to a coolant and how steep a temperature gradient you can sustain across it.

The chip vendor hands you two numbers that anchor the budget — the selected mode’s temperature limit and heat load — inside a wider thermal and mechanical envelope that also fixes supported interfaces and heat distribution. Tjunction-max — the maximum allowable on-die temperature, specified for the selected component and operating mode — bounds the qualified operating envelope; the component’s protection logic determines throttling or shutdown, and operation beyond its absolute limits risks damage. TDP — the thermal design power you must remove — is set by the silicon and the workload. Everything between the junction and the coolant is the resistance stack you must close, within the qualified package, TIM, plate, fluid and flow, and it is governed by a simple, unforgiving relation: the temperature rise from coolant to junction equals the heat removed times the thermal resistance of the path, ΔT = Q × Rθ. Fix Q (the TDP) and Tjunction-max, and the resistance you can afford collapses to a fixed budget. Spend it badly and the chip is illegal at any coolant temperature you can practically supply.

The resistance stack is a series chain — Tjunction to Tcase to the cooling medium — and like any series circuit, the largest resistor dominates. Walk it from the silicon outward, naming the reference temperature beside each resistance. OCP's cold-plate qualification method uses case-to-liquid-inlet resistance: do not add the full coolant rise again to an inlet-referenced result.

The junction-to-coolant thermal-resistance stack
StageInterfaceWhat it isWhy it dominates or doesn't
Junction → caseRθ-JC (in-package)Silicon → TIM1 → integrated heat spreader / lidLargely fixed by the vendor's package; you cannot improve it from outside
Case → cold plateTIM2 / thermal interfaceLid → second thermal interface → cold-plate basePoor contact or TIM degradation raises case-to-plate resistance; verify the assembled interface at the specified clamp load
Cold plate → coolantConvective resistanceMicrochannel / skived-fin base → flowing coolant filmSet by flow rate, channel geometry, and coolant; where DLC wins over air
Separate transport and exchanger quantitiesLoop ΔT; terminal HX approachesTechnology-cooling loop carries heat to the CDU heat exchangerClose the four-port energy balance separately; these are not additional solid resistors in an inlet-referenced component model
Keep the junction, case and coolant reference temperatures explicit. The component resistance model and the fluid transport/exchanger balances answer different questions; the worked decision below supplies the assumed resistance values.

The reason air loses at sufficiently high component heat flux is visible in the third row. The convective resistance from a surface to a fluid scales with the fluid's heat-transfer coefficient and the wetted area. Water's volumetric heat capacity is roughly 3,500× that of air, and its convective coefficient at a cold-plate surface is one to two orders of magnitude higher than forced air over a finned heatsink. Air can be pushed harder — more CFM, taller fins, colder supply — but each lever has sharply diminishing returns and a fan-power, noise and pressure penalty that can make further airflow uneconomic. Liquid operates in a different regime rather than merely beating air: it shrinks the case-to-coolant resistor by enough that the same TDP fits inside the same junction budget at a far more relaxed coolant temperature. → the cold-plate engineering is in Chapter 5.4; in-chip microchannels that attack Rθ-JC itself are in Chapter 16.2.

Why air hit a wall

Rack cooling is selected across overlapping equipment and facility envelopes; no universal rack-kW value is a physics cliff. Where earlier Parts speak of a “cooling cliff” or a “density wall”, read both as shorthand for the crossing engineered here — the point where a named rack’s heat split, flux and airflow limits stop closing against the hall you have — and not as a kW threshold. The cited source instead reports 30–40 kW as a typical RDHx range and more than 50 kW with active rear-door fans; those are door-reference conditions, while actual air and liquid-assisted systems close or fail at different duties according to rack airflow and inlet limits, system pressure, containment and recirculation, fan state, acoustics, design-day rejection, service access, and refresh profile. Any published rack, door, or cooling-unit capacity must therefore carry its stated air and water temperatures, flow, pressure, fan state, containment, and heat-capture conditions.

Three physical constraints explain why air becomes unattractive as heat flux and airflow demand rise. First, on a given system curve, fan power rises approximately with the cube of airflow: increasing CFM can impose a steep fan-energy penalty, and the result must be modeled at the actual pressure and fan-efficiency point. Second, air's low heat capacity forces large temperature rises and large volumes; the supply-to-return delta-T air can carry is small, and you run out of mass flow before you run out of fans. Third, acoustic and velocity limits cap how hard you can blow before noise, vibration, and bypass airflow make the hall unworkable and the cooling ineffective at the chip. These constraints make conventional room air unsuitable for the cited 132 kW rack, whose OEM record already assigns ~115 kW to liquid and ~17 kW to air. That product-specific heat split — not subtraction from a universal rack-kW ceiling — establishes its DLC requirement.

Where the supported equipment roadmap may require water at the rack, reserve the hard-to-retrofit structure, routes, isolation, CDU space, and rejection capacity early, even though air, RDHx, hybrid, and DLC envelopes overlap. An inherited air hall may or may not lack those provisions; inspect it rather than assuming. Adding missing liquid and structural interfaces later can run ~$2M/MW (cooling-only) to ~$5–6M/MW and up and still strand power or floor capacity. This is a named equipment-roadmap and facility-envelope decision, not a workload or rack-kW lookup. → retrofit paths are engineered in Chapter 5.10; air-system qualification is in Chapter 5.2.

The density curve, 2020–2027

The density wall would be an academic curiosity if accelerators had stayed put. Per-GPU thermal design power has climbed from the A100's ~300 W to the H100's 700 W to GB200's ~1.0–1.2 kW, with the standard Rubin (VR200) package at ~1.8 kW under Max Q and ~2.3 kW under Max P, while Rubin Ultra has no published per-package TDP. Multiply by the GPUs packed into a rack and the rack-level curve is steeper still, because the scale-up domain grew at the same time, concentrating more silicon behind a single liquid manifold. Across cited reference designs, the transition from H100-class air-capable configurations to the GB200 NVL72's declared liquid/residual heat split narrowed the feasible set from qualified air or RDHx paths to a supported DLC-plus-residual-air architecture.

Rack density by GPU generation, and the cooling regime each forces
Generation (year)Per-GPU TDPRack operating profiles / facility basisCooling regime forcedRelation to the air cliff
A100 / HGX (2020–22)~300–400 W~10–20 kWAir; raised floor + containmentComfortably under the wall
H100 / HGX (2023)~700 W~30–40 kWAir at the limit; RDHx optionalAt the wall; air still wins for many
GB200 NVL72 (2024–25)~1.0–1.2 kW132 kW nominal (HPE)Direct-to-chip liquid mandatory~3× over the wall; no air path exists
GB300 NVL72 (2025)~1.4 kW class135 kW TDP / up to 155 kW peak / up to 142 kW facility basisDLC; residual air load on RDHxWell over; hybrid liquid+air per rack
Vera Rubin NVL72 (2026)~1.8 kW Max Q / ~2.3 kW Max P188 kW Max Q / 228 kW Max P / 330 kW facility design basisDLC + 800 VDC power pathFar over; warm-water loops to free-cool
Rubin Ultra Kyber (2027)not published~600 kWDLC mandatory; in-chip microfluidics on the roadmapAn order of magnitude over the wall
The regime column is the consequence the density forces — it follows from the named rack's heat split and heat flux, not from rack kW alone.

The rightmost column is the consequence. As density and heat flux rise, the equipment profile narrows the feasible set, but selection still depends on liquid heat fraction, airflow/inlet limits, water availability, TCS/FWS conditions, climate/rejection, residual room heat, serviceability, redundancy, and the future tail. The decisions are when you cross for a named, funded refresh generation and whether the irreversible substrate is ready when you do; reserve infrastructure for those configurations and price later generations as options. What a hall scoped for a previous density cannot make up later is the substrate, not the fit-out — not the floor, not the power chain, not the cooling plant — which is why Chapter 1.1 prices the named ramp scenarios against that substrate before the IT is ordered.

30–40 kW typical RDHx; >50 kW with active fans; not a universal limit
SemiAnalysis/nVent cited RDHx range: 30–40 kW typical and >50 kW with active rear-door fans; verify the named door/rack and facility envelope
Scope & caveats

Select on the named door/rack, air and water conditions, fan state, containment, heat-capture target, residual room heat, climate/rejection, serviceability, redundancy, and future density.

Reference capacity, not a universal ceiling; verify named door/rack, water and air conditions, fan state, containment, and capture target.

~3,500×
volumetric heat capacity of water vs air — the reason liquid operates in a different cooling regime
132 kW nominal TDP: 115 kW liquid + 17 kW air
NVIDIA GB200 NVL72 by HPE: 132 kW nominal rack TDP, with 115 kW liquid and 17 kW air heat-removal duties
Scope & caveats

Exact HPE product profile. Keep this 132 kW / 115 kW liquid / 17 kW air record separate from the OCP MGX Rev. 1.1 reference profile of 120 kW / approximately 102 kW liquid / 18 kW air.

45 °C maximum liquid inlet; 65 °C maximum liquid return (separate limits)
QCT GB200 NVL72 QoolRack reference maxima: 45 °C liquid inlet and 65 °C liquid return; separate limits, not a selected operating pair
Scope & caveats

Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.

Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.

135 kW TDP; 155 kW peak; up to 142 kW facility design basis; ~90% liquid / ~10% air
per GB300 NVL72 rack (up to ~155 kW peak); CPUs/GPUs/NVSwitch liquid, optics/storage air
Scope & caveats

Size irreversible infrastructure to the facility design basis; run energy and TCO models on the operating profile.

~600 kWforecast
per Rubin Ultra Kyber rack (NVL144) on 800 VDC (roadmap, 2H2027)
Scope & caveats

NVIDIA's published figure (GTC 2025) is 600 kW per Rubin Ultra Kyber rack and GTC 2026 did not revise it. SemiAnalysis (2026-05-26) reports Kyber Ultra 'approaching 660 kW' — a single-source analyst estimate for a 2027 part, recorded here rather than adopted, since the vendor primary figure still stands.

~1.8 kW Max Q / 2.3 kW Max Pderived
Vera Rubin GPU package power — ~1.8 kW at Max Q / 2.3 kW at Max P
Scope & caveats

The same dual-die Rubin GPU package has two operating profiles; neither figure is a Rubin Ultra TDP. Rubin Ultra per-package TDP is not published.

~$5–6M/MW → greenfield parity
cost to retrofit an air-cooled hall across the cliff to AI liquid cooling; still strands capacity
Scope & caveats

20–40 kW/rack conversion targets; 100 kW+ conversions approach or exceed greenfield cost

~$2M/MW
cooling-only liquid retrofit, power-suitable hall — the ~80%-cheaper scope the full-conversion figure is often confused with
Scope & caveats

cooling scope only, structurally and electrically suitable hall

33% → 53% → ~60%
liquid-cooling penetration among AI chips — the installed-base adoption curve behind the density ladder
even in 2026 nearly half of AI chips still run on air — the two-speed market is real, and the liquid share is compounding fast

The cooling hierarchy and where each rung saturates

Map the density curve and required thermal interfaces onto the available cooling technologies and you get an overlapping ladder. Air, rear doors, cold plates and immersion each hand the load onward when their pressure, temperature or service envelope runs out. Choose the lowest-cost supported rung that clears peak duty in normal and degraded operation, because every added interface costs capital and service effort you do not get back. Added liquid plumbing spends integration effort, while retaining air spends fan pressure, floor area and inlet margin.

  • Air (containment + CRAH/in-row). Saturates when the selected rack’s airflow, pressure and inlet limits run out — a configuration limit, not a universal kW/rack cliff. No facility water at the rack; the simplest service path when the air system closes delivery and rejection at the lowest cost. Still the right answer for air-qualified storage, networking, HGX and RTX PRO nodes, and edge equipment. An NVL72-class rack that requires DLC needs its liquid path for inference as well as training: the rung follows the named rack profile, not the workload label. → Chapter 5.2.
  • Rear-door heat exchangers / air-assisted liquid. Bridges the air-cooled rack’s heat into water: passive doors spend server-fan pressure, active doors add their own fans to carry more coil pressure loss. Capacity belongs to the named door’s air/water schedule, not a universal passive or active band. Captures heat at the rack exhaust with a liquid coil; the brownfield-friendly rung when a water-fed door fits the existing server airflow, or a cold-plate L2A unit and room plant carry the full heat without building water. Saturates where the door coil can no longer extract a high enough fraction of the heat. → Chapter 5.3.
  • Direct-to-chip liquid (single-phase DLC). A mainstream 2026 path; one reported forecast is ~55%, but PMR's published category is broader cold-plate liquid cooling rather than measured single-phase-DTC deployment share. Cold plates on the GPUs/CPUs/switches, in-rack manifolds, dripless quick-disconnects, a CDU isolating the technology-cooling loop from facility water. Carries the cold-plate path into the 100 kW to 200+ kW rack-design problem when the selected rack, fluid and distribution close; extending warm-water loops to a Kyber-generation roadmap requires that generation’s own heat, pressure and service qualification. → Chapter 5.4; CDU loop in Chapter 5.6.
  • Immersion (single- and two-phase). Single-phase keeps the bath below boiling and owes you a wet-server service process; two-phase boils the dielectric and owes you vapor recovery as well as that process. PFAS fluid exposure remains a procurement issue where the named chemistry falls within the applicable definition. Compare auxiliary energy, wet support reactions, fluid recovery and warranty at matched duty before paying for either bath. → Chapter 5.5.
  • In-chip / direct-to-silicon microfluidics. The next rung, attacking the in-package Rθ-JC resistor itself with microchannels etched into or onto the die — the only lever that touches the dominant resistance the cold plate cannot reach. Roadmap, not yet default. → Chapter 16.2.
Deep dive: the chain of delta-Ts, approach temperature, and why warm water decides free cooling

Heat does not teleport from the junction to the sky; it walks down a staircase of temperature drops, and every step costs you. SemiAnalysis frames this as the Four Delta-Ts, and it is the right mental model for the entire facility. Start at the junction (~90 °C limit). Drop across the package and TIMs to the cold-plate coolant. The coolant warms across the IT load and returns to the CDU; account for that warming through mass flow, separately from any inlet-referenced resistance. Drop a third time across the CDU heat exchanger from the technology-cooling loop to the facility-water loop. Drop a fourth time at heat rejection — the cooling tower, dry cooler, or chiller that finally hands the heat to ambient. The junction temperature is fixed; ambient is fixed by your climate and season; everything in between is a budget of degrees you allocate across the component resistance, transport and exchanger boundaries.

The lever at each exchanger is approach temperature — the terminal temperature difference at one end of the exchanger — the stream leaving on one side against the stream entering on the other, a gap that never closes. A tighter approach means a more effective (and larger, costlier) exchanger but a warmer allowable FWS inlet for the same TCS supply. Formally this is captured by effectiveness and the NTU (number of transfer units) method: effectiveness is the actual heat transferred divided by the thermodynamic maximum, and it rises with NTU, which rises with exchanger surface area and overall conductance. More area buys more effectiveness buys a tighter approach buys warmer facility water for the same junction temperature.

Warmer water matters because it is the difference between free cooling and mechanical chilling. If your facility-water loop can run warm — ASHRAE's W17 through W45-plus classes key cooling water by supply temperature — a dry cooler or tower can reject heat to ambient for most or all of the year, and a low PUE becomes reachable when the annual plant curves and weather bins show that avoided compressor energy exceeds the added fan and pump input. If your delta-T budget forces cold supply water, you burn compressor energy on chillers, drive PUE up, and shrink your siting envelope to cool climates. Every degree you waste on a sloppy TIM or an under-sized exchanger upstream is a degree you cannot spend on free cooling downstream. This is why the 30 °C-coolant roadmap exists, and why warm-water design is treated as a first-class objective rather than an afterthought. → facility loops and warm-water design in Chapter 5.7; heat rejection in Chapter 5.8; the metric definitions in Chapter 15.1.

Thermal metrics used in this part

Part 5 leans on a small, consistent vocabulary of thermal metrics. Pin them down here so the engineering chapters can use them without re-deriving:

  • Approach temperature — the terminal temperature difference at one end of a counterflow exchanger: the stream leaving on one side measured against the stream entering on the other at that same end. For a CDU running TCS 45 °C in / 35 °C out against FWS 30 °C in / 40 °C out, the cold-end approach is 35 − 30 = 5 K; comparing the two outlets gives −5 K and describes the wrong relationship. Smaller approach, more effective and more expensive exchanger, warmer facility water that still serves the rack. The single knob you tune at every exchanger in the chain.
  • Effectiveness (ε) and NTU — effectiveness is actual heat transfer over the thermodynamic maximum; NTU is the dimensionless measure of exchanger size (conductance × area over the minimum heat-capacity rate). The ε-NTU method is how you size CDU and facility heat exchangers without solving the full temperature field. Higher NTU asymptotes toward ε = 1 with diminishing returns.
  • Delta-T (ΔT) — the temperature rise a coolant carries across a load. A larger ΔT moves the same heat at lower flow (smaller pumps, smaller pipes), which is why warm-water, high-ΔT design is favored. The flow-rate rule of thumb — ~1.25–2.0 L/min per kW — falls directly out of a roughly 7–12 K rise for water near 1 kg/L and 4.18 kJ/(kg·K). Use the actual fluid density and specific heat for selection, then check the component minimum flow and pressure limits.
  • Heat flux (W/cm²) — power per die area, the quantity the cold plate actually fights at the hotspot, distinct from total TDP.

The facility-efficiency metrics — PUE, WUE, ITUE, and TUE — sit one level up, scoring the whole plant rather than a single exchanger. PUE is total facility energy over IT energy; WUE is water consumed per IT energy; ITUE and TUE extend the accounting to capture fan and pump parasitics that liquid cooling reshuffles. These are defined canonically in Chapter 15.1 and used throughout Part 5 as the scorecard for the design choices this chapter sets up; we name them here only so the cross-references resolve.

Can four racks close the heat balance at both exchanger ends?

Trace. IT heat = 4 × 100 = 400 kW; liquid duty = 4 × 80.0 = 320 kW; air duty = 4 × 20.0 = 80.0 kW. Each water stream needs mass flow 320/(4.18 × 10.0) = 7.66 kg/s, or 459 L/min for the zone and 115 L/min per rack. Both heat-capacity rates are 32.0 kW/K. Cold-end approach = 35.0 − 30.0 = 5.0 K; hot-end approach = 45.0 − 40.0 = 5.0 K. Effectiveness = 320/[32.0 × (45.0 − 30.0)] = 0.667. Comparing the two outlets would test the wrong pair.

Component temperature = 35.0 °C + 1,000 W × (0.0150 + 0.0250) K/W = 75.0 °C. Margin is 85.0 − 75.0 = 10.0 K, so the specified 5.0 K reserve remains. Select this operating point for the thermal design and carry the separate 80.0 kW air duty into room cooling; omitting that branch strands equipment outside the cold plates.

Flip. At unchanged resistance and component heat, the reserve is exhausted at 85.0 − 5.0 − 40.0 = 40.0 °C inlet; warmer inlet fails. At fixed TCS temperatures, raising FWS inlet above 30.0 °C also fails the assumed 5.0 K cold-end approach requirement. A different exchanger selection or lower heat load must close the new point. OCP cold-plate qualification §1.3.2.1 defines the inlet reference; the balance remains canonical here. Hand heat and flow to 5.6, external head to 5.13, and installed acceptance to 13.5.

This chapter sets the foundation the rest of Part 5 builds on. Air pushed to its honest limit is Chapter 5.2; the rear-door bridge is Chapter 5.3; direct-to-chip liquid — the 2026 default — is Chapter 5.4; immersion is Chapter 5.5; the CDU and secondary loop are Chapter 5.6; facility water and warm-water design are Chapter 5.7; heat rejection is Chapter 5.8; and retrofitting across the cliff is Chapter 5.10. The density-and-cooling fork that this chapter treats as physics is framed as a scoping decision in Chapter 1.1; in-chip microfluidics that attack the in-package resistance live on the roadmap in Chapter 16.2; and the efficiency metrics that score every cooling choice are defined in Chapter 15.1.

Select the heat path only after the component margin, liquid/air split and both exchanger terminals close at the required load. A warmer operating point spends temperature reserve for rejection opportunity; a colder one spends plant energy for reserve. Carry the selected point into the fluid and hydraulic designs before buying equipment.

Cite this chapter
Fehn, J. (2026). Thermal Fundamentals & the Density Wall (Chapter 5.1). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-1-thermal-fundamentals-and-the-density-wall (accessed 2026-09-29).
@misc{aidc-5-1,
  author       = {Fehn, Jacob},
  title        = {Thermal Fundamentals & the Density Wall (Chapter 5.1)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-1-thermal-fundamentals-and-the-density-wall},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit