The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Cooling & Thermal Management › 5.12

Chapter 5.12

In this chapter · 7 sections
Term help

Cooling-Controls Transient Dynamics & Setpoint Stability

A direct-to-chip loop faces two transients — a heat step buffered by circulating inventory, and a flow loss that leaves the junction relying on locally coupled mass. Tune the controls against the first and qualify local protection against the second; only the actual heat imbalance and equipment envelope tell you how many seconds remain.

POWER-BOUNDGOODPUT

What you'll decide here

  1. How the CDU handles a maintained-flow heat step without hunting, where that tuning lives — CDU PLC, facility BMS or predictive supervisor — and how pump/CDU redundancy handles the separate flow-loss case; measure delays and stability against each equipment-qualified response window.
  2. Whether you ride a synchronized power slam with stored thermal/electrical inertia (BBU/BESS power-smoothing, a warm-water buffer volume, a flywheel-like reserve) or with a faster control response — and which is cheaper for your density tier.
  3. The dew-point margin and the condensation-risk policy on a rapid load drop: how many degrees above space dew point you hold secondary supply, and what the controls do when load collapses faster than the setpoint can rise.
  4. How cooling controls coordinate with GPU power-capping and the chip-to-BBU-to-BESS electrical spine so that a thermal excursion throttles compute gracefully instead of tripping it — the goodput-vs-trip decision.
  5. What you commission and demonstrate at Level 5: a real synchronized load-step test, an anti-hunting margin, and a dew-point excursion test — not just a steady-state capacity proof.

Chapter 4.5 told the electrical side of this story: a frontier training cluster is the worst load the grid has ever been asked to serve, because tens of thousands of synchronized GPUs ramp from idle to full power and back on millisecond timescales, and the power chain has to absorb the di/dt with stored energy before it propagates upstream. This chapter is the thermal twin of that problem. The same synchronized slam that stresses the busbar also lands, a few hundred milliseconds later, as a step change in heat flux at the cold plate — and a direct-to-chip loop may have far less usable thermal inertia than the legacy chilled-water reserve it replaces; only local component mass, circulating inventory and temperature margin establish that comparison. What this chapter forces is a controls decision rather than a capacity one — Chapter 5.1 owns the capacity wall — namely how the controls behave in the seconds around a transient: how fast they may move, how they avoid oscillating, and what they do when the load disappears as fast as it arrived.

The chapter turns on four settings: slew-limit tuning, inertia-vs-response, dew-point margin, and capping coordination. Set any of them wrong and the cost is concrete — a throttled GPU (lost goodput), a hunting loop (mechanical wear and unstable die temperatures), or condensation on a cold manifold inside an energized rack (the failure mode that ends up in Appendix F). None of this is exotic; it is the ordinary consequence of running a low-inertia loop under a load profile that swings harder and faster than any cooling plant was historically asked to follow.

The two transients: a heat step at flow, and a loss of flow

A legacy air-cooled hall looks like a giant thermal flywheel: room air, the raised-floor plenum, slab and racks, plus stored water in tanks and pipework. It rides through a load step only on the portion of that mass that couples to the heat during the event. A large room cannot by itself excuse a sluggish CRAH fan; establish excess-energy input, participating inventory, continued circulation and permitted temperature rise before assigning the delay budget.

A direct-to-chip loop can buffer a maintained-flow heat step for tens of seconds in its coolant and metal mass, but it offers almost no inertia after flow stops. The coolant in contact with the die is a thin film in a microchannel cold plate; the TCS inventory must come from the actual circuit schedule, and the transport time from cold plate to CDU heat exchanger and back is seconds, not minutes. The thermal resistance from junction to coolant is deliberately tiny — that is the entire point of liquid cooling — which means the die rides a small, power-dependent offset above the local coolant temperature. When GPU power steps by more than half of TDP in milliseconds (measured: power variations exceeding 50% of TDP within a few AC cycles; current swings from a 5–10 A idle baseline to 20–25 A in under 200 ms on a single inference node), it is that offset that jumps first — in milliseconds, buffered by the silicon and cold-plate mass — while the loop responds on the tens-of-seconds scale: the usable inventory and excess-energy integral below determine the loop excursion. At maintained flow the first loop-level symptom is the return temperature climbing; supply stays wherever the CDU holds it. A different boundary is loss of flow. Stop circulation and only the locally coupled fluid and solid mass can absorb heat on the immediate timescale — whether the junction throttles, shuts down or trips, and on what timescale, is product- and implementation-specific and must be established from OEM data, controls tests or an engineering calculation. That one is a redundancy problem (Chapter 5.11), not a controls-tuning one, and conflating the two produces controls designed against the wrong time constant.

This is the inversion that catches an operator migrating from air: a delay the old hall survived can throttle the new rack when local heat rises faster than the measured control and transport response. Establish those timescales; neither the air nor liquid label establishes their order. The HPE GB200 NVL72 loop is unforgiving precisely because the coupling is tight — ~165–236 L/min of secondary-loop flow per rack on the guide's water-property heat balance at a 7–10 °C design rise; Clariant PG25 gives ~172–246 L/min at the same rise, roughly 4% more (vendor reference designs run roughly 12 °C rise and specify ~110–140 L/min for the same ~115 kW liquid share — flow and delta-T trade directly, so never quote one without the other), and a die temperature that sits a small, power-dependent offset above supply (the rack accepts supply up to ~45 °C; the closer you run to it, the less headroom you hold) — and a loop that lags a load step costs you throttled GPUs. The control system now lives in the goodput path.

CDU slew limits and the lag-vs-stability fork

The CDU is where the control decision physically happens. Its two actuators are the pump VFD (setting flow) and a modulating control valve — usually a three-way mixing valve on the facility-water side of the heat exchanger, which sets how much heat the TCS loop rejects and therefore the secondary supply temperature. Both actuators have slew limits: the maximum rate at which they are permitted to change. Specify the CDU manufacturer’s allowed pump-RPM change, valve travel per minute, flow, pressure and steady-state supply-temperature band. Those are the slew and control limits you tune against the actual transport delay and duty response; a predictive-control experiment on one vendor’s unit does not set another unit’s actuator limits.

The slew limit cuts both ways. Set it too slow and the loop falls behind a sustained load step: the return comes back hotter, the heat exchanger rejects less than the racks are producing, and over tens of seconds the accumulated energy drags secondary supply above setpoint until the die's headroom is spent and the GPUs throttle — lost goodput, the exact thing a frontier training run cannot afford. Set it too fast and you invite the opposite failure: the valve and pump overshoot the correction, the over-corrected temperature transports back to the sensor a few seconds later, the controller over-corrects the other way, and the loop hunts — a sustained oscillation in flow, temperature, and pressure that wears valve seats and pump bearings, fatigues quick-disconnects and hoses with pressure cycling, and swings die temperatures around the throttle threshold so the GPUs themselves oscillate between full and capped. Hunting is the cooling-controls failure mode catalogued in Appendix F, and diagnosing it means distinguishing tuning faults such as excessive gain and transport delay from hardware faults such as valve stiction or a failed sensor, as well as unstable hydraulic authority.

The reason this is hard is the transport lag in the callout above: any loop with significant dead time between actuation and feedback is prone to oscillation if its gain is too high. Classical fixes apply — lower proportional gain, lengthen the integral time, add a setpoint deadband so small excursions do not chase the actuator, and enable anti-windup so the integrator does not saturate while the valve is slew-limited at its travel cap. The next lever is to stop relying on feedback alone where timely feedforward demonstrably improves the response enough to repay its telemetry and failure dependencies.

Three control strategies for a low-inertia loop under synchronized slams
StrategyHow it answers a slamThrottle riskAdded capital / complexityBest fit
Reactive feedback (PID only)Senses temperature rise, then slews valve/pump within limitsMeasure excursion against the actual marginLowest; native CDU PLCAccept when the demonstrated response meets the load and fault envelope
Feedforward / predictive supervisorReads the GPU power/utilization signal and pre-positions valve and pump before the heat arrivesVerify improvement and stale-signal fallbackTelemetry integration, supervision and requalificationUse when repeatable load information improves the tested response
Absorb with thermal inertia (buffer)Rides the heat step on stored thermal energy in dedicated warm-water buffer volumeDepends on maintained-flow participationHighest; buffer tankage sized for thermal ride-throughUse when the excess-energy integral exceeds the existing inventory margin
The three strategies are feedback-only reactive control, feedforward/predictive, and absorb-with-inertia. Figures are 2026 practitioner/vendor ranges; see keynumbers for sources and vintages. Each row trades response speed, capital, and stability differently.

The three strategies are a spectrum, and real facilities blend them. For dense training clusters whose load steps are large, fast, and — critically — synchronized and predictable, feedforward answers a timing problem by acting before delayed temperature feedback arrives, when its load signal is timely and the measured response beats the qualified feedback-only case. A training step is a repeating waveform: the cluster slams to full power during compute, drops during a collective or a checkpoint, slams again. If the controls can see that waveform — via the GPU power-management telemetry, the scheduler, or even the BMC power signal — they can feedforward the valve and pump position so the coolant is already cold when the heat arrives, instead of chasing it. In ProphetStor's disclosed integration test, a 20-job mixed workload on a single eight-GPU Supermicro server, mapping coolant flow to measured GPU power cut pump energy 16–18%. The supplier's separate 22–28% figure for total CDU energy is not established as a result of that test. The bottom row — buy your way out with inertia — is the fallback when even feedforward cannot keep up, and it is where the cooling and electrical problems converge.

The electrical spine: cap thermal load, smooth upstream draw

The cleanest way to make a cooling-controls problem tractable is to reduce the heat step at its source; GPU power-capping does that, while the BBU/BESS spine in Chapter 4.5 smooths the upstream electrical boundary without changing chip heat. Coordinate the levers by what each actually changes:

GPU power-capping is the fastest and cheapest. The accelerator's own power-management firmware can clamp or ramp-limit board power, smoothing the di/dt at the source. NVIDIA’s 2026 Vera Rubin MGX design puts a few hundred joules of storage per GPU at the rack boundary: its published 400 J/GPU capacitor figure describes electrical power smoothing, not on-die thermal storage. Capacitors, BBU and BESS can flatten upstream draw while unchanged silicon power still makes unchanged heat; Chapter 4.5 owns their location and duty. A cap is also the graceful-degradation lever for cooling: when a thermal excursion is developing, throttling compute a few percent to hold die temperature is almost always cheaper than letting the loop chase it into instability or, worse, letting a chip trip. This is the goodput decision — a small, deliberate cap costs a sliver of throughput; an uncontrolled excursion costs a restart.

Battery-backed units (BBU) and facility BESS absorb the electrical transient so the grid sees a smoother draw. NVIDIA's production-BESS guidance for AI factories names transient absorption and power-smoothing as first-class BESS roles alongside ride-through, with closed-loop state-of-charge control sized to the workload's transient signature. Be precise about what that buys thermally: storage changes where the electrons come from, not how much heat the silicon makes — the rack's heat flux follows GPU power exactly, wherever the power is sourced, so only the power cap above actually flattens the thermal waveform. What the BESS buys the cooling plant is indirect but real: it keeps upstream electrical disturbances from becoming compute disturbances, so the thermal profile stays the workload's own.

The coordination decision is who arbitrates. If GPU capping, BBU smoothing, and CDU controls each react independently to the same slam, they can fight: the cap reduces heat just as the CDU finishes slewing colder, and now the loop overshoots cold (toward the dew-point floor, below). The mature pattern is a supervisory layer that knows the load forecast and sequences the responses — cap and smooth first, feedforward the cooling to the post-cap heat profile — rather than three reactive loops racing each other. → Chapter 4.5 owns the electrical-spine design; Chapter 14.7 owns the in-operation power-and-thermal management policy.

How much condensation margin does this error budget require?

Minimum commanded supply = 20.0 + 0.5 + 0.2 + 0.8 + 0.5 = 22.0 °C. With that command the budget closes. If the local dew point rises while the command remains 22.0 °C, the crossover is 20.0 °C; above it the stated reserve fails, so raise the permitted command or invoke the qualified fallback before exposing colder surfaces. Verify the coldest surface and startup state, not only the CDU sensor. ASHRAE environmental guidance provides condensation context; the fluid/equipment envelope remains in 5.7 and installed tests in 13.5.

The dew-point floor interacts with the warm-water trend in an awkward way. Everything in Part 5 pushes coolant warmer — 30 °C-plus facility water to maximize free cooling and heat reuse (Chapter 5.1) — and warmer supply is comfortably above any realistic dew point, so condensation risk falls as supply temperature rises. Good. But warmer supply also means a smaller temperature margin between coolant and the die's throttle limit, which makes the loop less forgiving of a transient overshoot: there is less thermal headroom to absorb a momentary lag before the GPU throttles. So the warm-water decision is itself a transient-stability decision — it trades condensation risk down for throttle-margin risk up, and the control loop has to be tuned for the regime you actually run, near the middle of the operating band, not at a corner.

>50% TDP
GPU power swing within a few AC cycles (~40–80 ms); idle-to-full step under ~200 ms on an inference node
30–50 ms
PSU "inertia" lag between GPU current step and AC-input response; upstream converter bandwidth only a few kHz
±3% RPM/min, ≤10% valve/min
CDU pump and valve slew limits reported in one vendor's predictive deployment
Scope & caveats

Supplier-reported limits, not OCP requirements or a market-wide actuator envelope. The separate integration example uses one Supermicro server with eight H100 GPUs and operator-applied steady-flow recommendations.

Obtain the installed CDU/controller limits and validate the required transient locally.

22–28%
total CDU energy saved with predictive flow control, from the supplier's reported benefits; not established as the result of its separate 20-job, single-server test (pump energy alone: 16–18%)
Scope & caveats

The reported range belongs to the supplier’s Proven Benefits discussion; it is not established as the result of the separate twenty-job, single-server integration test.

A vendor-reported result at its stated boundary; compare matched useful output and facility energy before selecting predictive supervision.

±0.5 °C
secondary-supply temperature a well-tuned CDU holds at setpoint in steady state
~2 °C above dew point
minimum secondary-supply margin CDUs hold to prevent condensation on the TCS loop
~110–140 L/min (vendor references at ~12 °C rise)
GB200 NVL72 coolant flow in vendor reference designs at roughly 12 °C rise
Scope & caveats

The vendor-reference band is observed in QCT and nVent designs at roughly 12 °C rise. The guide's separate ~165–236 L/min water-property design-flow derivation uses a tighter 7–10 °C rise across ~115 kW; Clariant PG25 gives ~172–246 L/min at the same rise, roughly 4% more.

QCT ≤130 L/min and nVent ~138 L/min are vendor-reference operating points; using the separate water-property guide band requires the guide's 7–10 °C design rise, while Clariant PG25 at the same rise gives the slightly higher ~172–246 L/min band.

~400 J/GPU
rack-level power-smoothing energy targeted for the Rubin generation — flattening transients at the source
Scope & caveats

NVIDIA's stated design figure for the Vera Rubin power-shelf PSU capacitor system (NVIDIA Vera Rubin POD, 16 March 2026); the platform entered full production in August 2026. Vendor design statement, not an independent field measurement.

Anti-hunting: tuning the loop so it does not eat itself

Hunting is the most common preventable failure in a commissioned liquid plant, and it is invisible at steady state — it only shows up under the synchronized-load conditions the workload actually creates. The mechanism is dead-time-driven oscillation: a loop with several seconds of transport delay between actuator and sensor will oscillate if its loop gain is high enough that a correction returns to the sensor still large enough to provoke an opposite correction. Three knobs and one architecture decision keep it stable.

  • Gain and integral time. Lower proportional gain so a single correction does not overshoot; lengthen integral time so the controller does not wind up faster than the loop can respond. The classic symptom of too-aggressive tuning is a regular swing in supply temperature, often roughly sinusoidal; compare synchronized actuator and sensor traces to identify its period and cause instead of assigning a universal multiple of transport delay.
  • Deadband. A small setpoint deadband (e.g. ±0.3–0.5 °C where measured noise and allowable excursion justify that band) stops the loop from chasing sensor noise and micro-excursions, which is what drives high-frequency valve dithering and seat wear. The cost is a slightly looser steady-state band — an easy trade.
  • Anti-windup. Because the actuators are slew-limited, the integrator will try to demand more correction than the valve can deliver during a fast transient. Without anti-windup the integrator saturates, and when the transient passes the loop dumps the accumulated error as a giant overshoot — straight toward the dew-point floor on a load drop. Anti-windup (clamping or back-calculation) is not optional on a slew-limited loop.
  • Architecture: feedforward can lead a dead-time loop. Feedback cannot erase physical transport lag; a timely load signal can feedforward the expected correction while the local PID trims residual model error. Accept that path when it improves the bounded response and falls back safely on stale data. This is the same insight as the predictive-supervisor row in the table — it is the difference between chasing the load and leading it.

One more anti-pattern worth naming: controls fighting across layers. If the CDU PLC, the facility BMS, and a GPU-side thermal manager each run their own loop against overlapping setpoints, they interact like coupled oscillators and the whole system hunts even when each loop is individually stable. Decide which layer owns the secondary-supply setpoint and let the others observe, not actuate. This is the thermal analogue of the electrical arbitration problem above.

Deep dive: sizing a warm-water buffer to ride a synchronized step

Required effective volume follows the time integral of load minus heat rejection, divided by the fluid’s volumetric heat capacity and allowed temperature rise. Credit existing volume only when maintained circulation, transit and mixing make it participate inside that interval. Nameplate tank volume is not the result. Source-level power capping changes the heat imbalance; BBU/BESS at unchanged IT power does not. The following case shows when the required additional tank disappears.

Does this heat step require additional buffer volume?

With no incremental rejection, excess energy = 60.0 × 10.0 = 600 kJ. Required effective volume = 600/(4.18 × 2.00) = 71.8 L; additional volume = max(0, 71.8 − 50.0) = 21.8 L. Select added storage only if the actual tank supplies that effective contribution within the interval; this result does not qualify lost circulation.

With linear rejection recovery, the excess-energy triangle is 0.5 × 60.0 × 10.0 = 300 kJ. Required volume is 35.9 L, so the existing 50.0 L suffices and added volume is zero. The crossover excess energy is 50.0 × 4.18 × 2.00 = 418 kJ; below it a new tank buys no required capacity. Demonstrate the imbalance and participation before choosing tankage over control changes. ASHRAE’s energy-balance boundary provides measurement context; 5.1 owns conservation and 13.5 accepts the installed transient.

Commissioning the transient, not just the capacity

The operational consequence is that a cooling plant which passes a steady-state capacity test can still fail catastrophically under the workload's real transient profile, because the two are different physics. A Level 4 capacity proof says the plant can reject the rated heat at the rated flow indefinitely. It says nothing about whether the loop hunts when the load steps, whether the controls catch a slam before the die throttles, or whether the setpoint clamps above dew point on a fast drop. Those have to be demonstrated explicitly — which is why the integrated-systems-test (IST) for a liquid-cooled hall has to include transient cases, not just the load-bank-at-full sequence.

The transient acceptance set is short and specific: a synchronized step test (drive the load from low to full and back as fast as the load banks or a real workload allow, and confirm the controls catch it without throttling and without overshoot), an anti-hunting demonstration (hold the load at the worst-case operating point and confirm the loop settles inside its band without sustained oscillation), and a dew-point excursion test (drop the load fast and confirm the setpoint clamp holds secondary supply above the measured dew point throughout). Pair these with the metering and telemetry that make them observable in operation — per-rack power comes off the electrical metering of Chapter 4.12; the flow, supply-temperature and dew-point signals are a thermal instrument schedule you specify here — sensor placement, accuracy, sample rate and timestamp alignment to the power stream — because no electrical meter delivers them. → cooling acceptance and CDU commissioning in Chapter 13.5; the full Level 5 IST and failure-mode demonstration in Chapter 13.6.

This chapter is the thermal twin of the electrical-transient problem in Chapter 4.5 (UPS/BBU/BESS ride-through and transient absorption), and the two must be designed together. The capacity wall and warm-water rationale that set the throttle and dew-point margins live in Chapter 5.1; the DLC loop architecture (TCS/FWS separation, CDU, flow and delta-T) in Chapter 5.4; thermal reliability, leak detection, and the consolidated coolant-leak/thermal-runaway catalog in Chapter 5.11 and Appendix F. The telemetry that makes transients observable is metered in Chapter 4.12; the goodput-vs-availability framing behind the capping decision in Chapter 12.2; cooling/CDU commissioning in Chapter 13.5 and Level-5 IST in Chapter 13.6; and in-operation capacity, power, and thermal management in Chapter 14.7.

Choose the control bandwidth and buffer from the actual delay and excess-energy history while preserving local limits. Additional inventory helps only while heat can reach it; predictive supervision earns its place through a matched improvement and loses authority when its inputs fail.

Cite this chapter
Fehn, J. (2026). Cooling-Controls Transient Dynamics & Setpoint Stability (Chapter 5.12). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-12-cooling-controls-transient-dynamics-and-setpoint-stability (accessed 2026-09-29).
@misc{aidc-5-12,
  author       = {Fehn, Jacob},
  title        = {Cooling-Controls Transient Dynamics & Setpoint Stability (Chapter 5.12)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-5-cooling-and-thermal-management/5-12-cooling-controls-transient-dynamics-and-setpoint-stability},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit