Chapter 14.7
In this chapter · 8 sections
Capacity, Power & Thermal Management in Operation
The megawatts you energized are a fixed, capital-intensive ceiling; steady-state operations trades how full you run that budget against how violently the workload can swing it.
What you'll decide here
- Whether to oversubscribe the power budget, and by how much — POLCA’s roughly 3% training and 21% inference headroom are measured workload cases, so qualify your correlated peaks and SLO before assuming the same policy avoids a feeder trip.
- What the rack-, row- and facility-level power-cap hierarchy is, who can change it and how fast the installed controls act — while independent protection settings and stored-energy ride-through keep the feeder inside its time-current envelope when BMC or agent control cannot respond.
- How much stranded capacity you will tolerate, and which of the three strands (cooling-limited, power-chain-limited, fragmentation-limited) you are willing to spend capital to recover versus simply schedule around.
- Whether the facility participates in demand response / curtailment, and which workloads are eligible — because the revenue is real but the wrong workload curtailed at the wrong moment destroys more goodput value than the demand-response payment is worth.
- What forecasting cadence and reserved-capacity buffer the capacity-planning function runs at, so that the density ramp (40 kW → 132 kW → 600 kW racks) never arrives faster than the power and cooling substrate can absorb it.
By the time a facility reaches steady-state operation, its single most expensive and least reversible decision — how many megawatts it can energize — is already made and paid for. The interconnection slot is won or lost, the transformers are humming, the cooling plant is sized. What remains is an operational discipline the chip-bound era never had to learn: running a fixed power budget hot enough to earn its capital back, without running it so hot that a workload transient trips it. This is the day-2 face of the power-bound thesis. The binding constraint has shifted from how many GPUs you can buy, or even cool, to how completely you can convert energized megawatts into goodput, request by request and step by step.
This chapter treats four operational functions as a single problem because they share one ceiling: live power management and oversubscription (running the budget full), transient management in operation (keeping the swings inside the budget), stranded-capacity and thermal operations (recovering megawatts the design left on the table), and capacity planning, workload-aware operations, and demand response (matching the ramp and the grid). The physics of the transients themselves is canonical in Chapter 4.5 (storage and ride-through) and Chapter 5.12 (cooling-controls dynamics); here we manage them operationally.
How full do you run the power budget?
A facility provisioned for 100 MW of IT load rarely draws 100 MW. Nameplate is the sum of every device's worst-case rating; real workloads almost never align their peaks. The gap between nameplate and realized peak is power headroom, and the central operational question is whether to sell it — to oversubscribe, deploying more racks than the budget would support at nameplate, betting that they will not all peak together. The bet can pay well: the POLCA study reported roughly 30% more servers under the same power budget with negligible performance loss in its evaluated inference setting; validate the result for your request mix and SLO. It can also fail: the budget overruns, the protection scheme fires, and a statistical near-miss becomes a real outage.
Whether to take the bet comes down to whether your workload leaves headroom to sell, and here the two archetypes diverge sharply. Within a tightly phased synchronous job, compute, collective and checkpoint phases can correlate demand; all-reduce does not universally coincide with peak GPU board power. That correlated swing leaves almost nothing to overbook — the POLCA study’s measured training headroom was on the order of ~3%, not a universal ceiling. With sufficiently dispersed request arrivals, compute-bound prefill and memory-bound decode can overlap at different phases across inference nodes, smoothing aggregate draw below the sum of peaks — the POLCA study reported headroom on the order of ~21% for its inference profile; correlated prefill or synchronized arrivals can erase that margin. The same overbooking ratio can add usable capacity under one measured profile and breach a feeder budget under another.
| Archetype | Peak correlation | Power headroom | Oversubscription posture | Backstop if budget overruns |
|---|---|---|---|---|
| Synchronous pre-training | Measure phase alignment and compute/communication overlap | ~3% | Minimal to none; size to measured correlated peak | Fast rack/row power cap → frequency throttle; checkpoint-tolerant of the slowdown |
| Post-training / RL (rollouts) | Mixed — async rollout pool smooth, trainer correlated | Between training and inference; disaggregate to measure separately | Oversubscribe the rollout pool; protect the trainer | Cap rollout pool first (interruption-tolerant); shield trainer |
| Online inference | Arrival mix and prefill/decode correlation determine peaks | ~21% | POLCA’s ~30% result applies to its evaluated profile; qualify SLO | Priority-based capping; throttle low-SLO tenants first |
| Batch inference | Batch launches can correlate; scheduler can shape admission | Measure headroom under the selected queue policy | Use measured margin and deadline slack | Pause/defer within deadlines; price replay and delay |
The table separates two tests: measured peak correlation sets headroom, and interruption/recovery cost sets the price of using it. Batch inference can absorb deferrable work when deadline slack exists; a synchronized batch launch can still create a sharp peak. Price the actual queue and waveform before assigning it oversubscription risk. Sophisticated operators do not pick one posture; they run a heterogeneous budget where the interruptible, uncorrelated workloads absorb the overbooking that the synchronous, latency-bound workloads cannot. → the disaggregated RL picture is in Chapter 1.4; the inference burst profile in Chapter 1.3.
The power-cap hierarchy: the budget's enforcement layer
Oversubscription is a bet, and every bet needs a stop-loss. In a power budget that stop-loss is the power-cap hierarchy — a layered set of enforcement points that manage realized draw within a validated workload envelope; independent electrical protection still covers faults and loss of control. The hierarchy spans four levels, each slower and broader than the one below it. At the GPU/node level, the accelerator’s hardware/firmware power-management path reduces clocks within its specified measurement and actuation envelope; a microsecond device response must be supported by the named platform, not inferred from a BMC command. At the rack/PDU level, intelligent rack PDUs and the BMC enforce a rack budget across nodes. At the row/lineup level, the DCIM control plane allocates a shared budget across racks fed from a common busway or RPP. At the facility level, the EMS holds the campus draw under the interconnection limit and any demand-response obligation.
The design question is where the binding cap sits and how fast it engages. Cap too low and you leave goodput on the table — GPUs throttled that the budget could have fed. Cap too high, or too slowly, and a correlated transient breaches the budget before the control loop catches it. The hard constraint is speed: participating ranks can produce a millisecond-scale step as compute resumes after a collective. Measure overlap, step size and the installed response deadline; supervisory polling cannot be credited with faster actuation than it demonstrates. Time-critical response therefore belongs in validated local device controls and the designed power/energy buffers, with independent protection coordinated to the equipment envelope. The rack BMC issues supervisory cap commands on an out-of-band loop measured in seconds, and the DCIM and EMS layers set budgets the fast layers enforce rather than reacting to transients themselves. Specify measurement delay, command latency, actuator response and behavior after control-plane loss for the platform you actually bought, and keep oversubscription inside the equipment's time-current and ride-through envelopes: feeder safety is a protection setting, never an assumed BMC reaction time. An EMS whose sensing and actuation exceed the transient deadline must leave that duty to the qualified local path.
Trace. Supervisory response bound = 1 s + 2 s = 3 s. Energy required = 4 MW × 3 s = 12 MJ, exceeding the 8 MJ usable buffer by 4 MJ. Maximum supported delay on this rectangular-step model = 8 MJ / 4 MW = 2 s. Select HOLD on the proposed oversubscription: retain the headroom until a local control/buffer design satisfies both the power and energy limits. A successful BMC API call is not proof of actuation, and a breaker trip is not normal budget control.
Flip. The energy test changes to pass only if the validated total delay is no more than 2 s, or usable energy is at least 12 MJ at the stated 3 s delay, with discharge-power capability still sufficient. A proposed local response bounded at 0.5 s would require 2 MJ, but it earns operating authority only after latency, stale-data, saturation and control-loss tests; time-current coordination and protective devices still constrain operation. The method is the energy integral ∫Pexcessdt and the SI watt–joule–second relationship. Chapter 4.5 owns storage duty, Chapter 4.11 protection, and this chapter owns whether the operating response fits both.
Transient management in operation
The power transient is what makes oversubscription dangerous in the first place. A tightly phased synchronous job can pulse: compute raises demand, while exposed communication or checkpoint phases can lower it. Overlap and staggered jobs reshape that waveform; a job start, kill or reload changes the participating load, not automatically the entire facility. At fleet scale workload-driven swings are enormous and abrupt, but the July 2024 Virginia case is not evidence of their amplitude. That six-fault, 82-second 230 kV sequence caused protection-driven loss of ~1.5 GW of data-center load; it is a grid-fault ride-through case that helped drive NERC's Level 3 alert (2026) and made ride-through a first-order planning input. The requirement is now becoming enforceable: FERC Order RD26-7-000 (2026-07-16) directs NERC to file mandatory computational-load Reliability Standards by year-end 2026, and ERCOT's NOGRR 282 put demand-side ride-through into its operating guide (effective 2026-08-01). A campus that sheds 100+ MW on a voltage sag is on its way to being a compliance event, not just a goodput event. The physics of absorbing these swings — UPS, BESS, ride-through, and the controls that damp them — is engineered in Chapter 4.5 and Chapter 5.12. The operational question this chapter owns is narrower and continuous: how do you run the live facility so the transients you generate stay inside what your absorption layer and your grid contract can take?
The answer is a set of operational levers that shape the workload's power profile before it reaches the power chain. Power smoothing deliberately holds a floor under idle phases (a synthetic GPU load during the all-reduce trough) so the swing the grid sees is smaller than the swing the GPUs actually produce — trading a little energy for a lot of transient amplitude. Ramp-rate limiting staggers job starts and checkpoint reloads so a 100 MW cluster does not step its entire load in one second; the scheduler launches ranks in waves. Stagger and phase-offset scheduling spreads correlated peaks across time on purpose. And the storage layer — UPS and BESS — absorbs whatever residual swing the workload levers leave. How much of the transient to suppress in software (cheap energy, free capital, but it taxes goodput and is the operator's job) versus in hardware (BESS sized to the swing — capital up front, no goodput tax) is the design decision; most facilities do both, and the ratio is where they differ.
Deep dive: why the all-reduce barrier is a power-systems event, not just a networking one
Gradient synchronization is easy to file under networking and forget at the switchboard — an operating blind spot distinct from the grid-fault ride-through failures addressed by the NERC Level 3 alert. Consider the mechanism. In a profile without compute/communication overlap, a synchronous data-parallel step can expose two different power signatures: a compute phase where every GPU runs forward/backward at or near TDP, and a communication phase — the all-reduce — where GPUs largely idle while the fabric exchanges gradients. Where the barrier tightly aligns participating ranks, the cluster's aggregate power is not the smooth average of thousands of independent devices; it can approach a square wave in a tightly phased profile, but overlapping compute/communication, job staggering and device control change both amplitude and shape. Measure the rack and upstream trace at sufficient bandwidth.
Now scale it. On a 100 MW training cluster that swing can be tens of megawatts, stepping in well under a second, thousands of times a day. To the GPU it is a clock-frequency artifact; to the upstream transformer, the UPS, and the utility feeder it is a load transient with real di/dt, voltage-sag, and frequency-excursion consequences. Measure those workload-driven transients directly: the July 2024 rise to 60.047 Hz and 1.07 pu followed protection-driven loss of ~1.5 GW during grid faults and does not quantify all-reduce phase coherence. POLCA’s roughly 3% training versus 21% inference headroom belongs to its measured profiles; neither a perfect training oscillator nor independent inference arrivals is a universal premise. The operational levers above — power smoothing to fill the trough, ramp-rate limiting to soften the edge, BESS to absorb the residual — are what determines whether the utility will interconnect the facility at all. → the absorption hardware in Chapter 4.5; the grid-side instantaneous-loss data in Chapter 4.5 and siting/interconnection in Chapter 3.2.
Stranded capacity: the megawatts the design left on the floor
Stranded capacity is energized, paid-for power or cooling that cannot be converted into compute because some other resource caps out first. In a power-bound world that is capital spent on megawatts that earn nothing. It comes in three distinct strands, and the diagnosis matters because the recovery for each is different — confusing them wastes the recovery capital too.
- Cooling-limited. The power chain can feed the racks but the cooling plant cannot remove the heat, so racks sit half-populated or throttled. This is the dominant strand in halls built for air and asked to host liquid-class density — the cooling cliff showing up as a utilization ceiling. Recovery is mechanical: more heat-rejection capacity, warmer water, or a cooling retrofit. → Chapter 5.4.
- Power-chain-limited. Cooling has headroom but a transformer, busway, RPP, or breaker is the binding constraint — often because the design reserved fault-margin or redundancy capacity that nameplate planning never released to load. Recovery is electrical and partly statistical: oversubscription releases reserved headroom; rebalancing phases and circuits recovers fragmented capacity. → Chapter 4.6.
- Fragmentation-limited. Both power and cooling have aggregate headroom, but it is scattered — a kilowatt free here, two there — in pockets too small to land a 132 kW rack. This is the most insidious strand because the DCIM dashboard shows spare capacity that the placement engine cannot use. Recovery is operational: workload-aware placement, defragmentation campaigns, and consolidating partial rows.
Misdiagnose the strand and you spend in the wrong place. Add chillers to a fragmentation problem and you have bought cooling you still cannot fill; oversubscribe a cooling-limited hall and you trip the thermal protection you were already brushing against. Reading which resource is actually binding, per row, from telemetry before committing recovery capital is the operational core of stranded-capacity management. → the telemetry that reveals the binding constraint is in Chapter 14.2.
Thermal operations: running the cooling plant against the load
Thermal operations is the cooling-side twin of power management: the continuous job of matching heat rejection to a load that swings as violently as the power draw does. The same all-reduce that pulses the power chain pulses the heat load — and the cooling plant's thermal mass and control loops respond on a slower timescale than the electrical ones, which is its own hazard. At maintained flow, coolant and metal mass may buffer a heat step, but a cooling-loop loss that stops flow is implementation-specific: the selected rack's throttle, controlled-shutdown, emergency-shutdown and no-response states require OEM transient data, controls tests or an engineering calculation. That evidence sets the required pump/rejection protection and redundancy topology.
The day-2 levers are setpoint and flow management across four separate limits: the facility-water supply class, the CDU approach, the selected rack's qualified TCS inlet/return temperatures, and its OEM flow/pressure-drop envelope; keep the secondary supply above white-space dew point to stay 100% sensible. The steady-state choice is how warm to run the loop. Warmer water expands free-cooling hours and crushes the cooling share of PUE (warm-water DLC plants approach ~1.1 versus ~1.3–1.5 for typical air halls, though well-optimized air can also reach ~1.1), but it shrinks the thermal margin to the throttle threshold — a GB200 that deviates from its inlet spec throttles, on a product- and workload-specific curve you measure rather than assume. Run cold and you waste compressor energy buying margin you may not need; run warm and you are operating closer to the cliff, where a transient that would have been absorbed becomes a throttle event and a goodput loss. Cooling-controls stability under these transients is canonical in Chapter 5.12; the piping and CDU mechanics in Chapter 5.13; predictive maintenance of the cooling plant in Chapter 14.5.
Scope & caveats
Observed headroom in the historical fleets POLCA studied (Patel et al., ASPLOS 2024); POLCA separately simulated 30% additional provisioned servers under specific controls. Neither figure is a deployable allowance for a 2026 reasoning, MoE, or disaggregated fleet: derive that from the proposed fleet's measured coincident demand, its tested protection and capping response, and the serving degradation you will allow.
Scope & caveats
POLCA oversubscription gain
Scope & caveats
Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.
The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.
Scope & caveats
First-order model over 22 balancing authorities using 2016–2024 load histories and constant added load. At 0.5% curtailed annual energy, mean event duration is 2.1 hours; 177 hours per year have some curtailment, distinct from about 44 full-load-equivalent hours. Local network capacity, ramping and ramp-feasible reserves are outside the model.
National modeled potential does not allocate service to a parcel; require the local power-flow/stability study and executable offer.
Scope & caveats
Treat the figures as illustrative facility benchmarks, not technology-wide ranges. Climate, economization, electrical design, utilization, and operating discipline can let air-cooled facilities reach PUE at or below 1.1.
Scope & caveats
Respondent-based survey result (Uptime Institute, Annual outage analysis 2026, Figure 4, n=96, 'primary cause of your most recent impactful incident'), not an event-by-event census of the industry. The 2026 report does republish 45% for 2025 outages — down from 54% in 2024 — and reports a fifth consecutive year of declining per-site outage frequency with the pace of improvement slowing. Uptime attributes part of the fall to electrical upgrades and part to other causes rising (fire-suppression share up six points year-over-year).
Capacity planning and forecasting under a hard ceiling
Capacity planning in a power-bound facility is the discipline of never letting demand for megawatts outrun the supply you have energized, while never leaving so much energized capacity idle that the capital decays unused. It is a forecasting problem with an asymmetric loss function: under-provision and you turn away revenue against a depreciation clock that runs whether the racks are full or not; over-provision and you have stranded energized megawatts, the costliest resource on the site. What the forecast must track is the density ramp, not just steady-state load. The substrate must be planned against the generation you will deploy next, because the irreversible parts (transformer capacity, busway ampacity, cooling-plant tonnage, floor loading) cannot be retrofitted on the timescale that rack density is climbing: 40 kW air-cooled racks gave way to ~132 kW NVL72, heading to Vera Rubin NVL72 (188 kW Max Q / 228 kW Max P; 330 kW facility design basis) and ~600 kW Kyber-class. A hall whose power and cooling were sized to today's density strands its own future.
The operational choice is the reserved-capacity buffer: how much energized headroom you hold back against the next ramp step and against failures. Hold too little and the next GPU generation arrives with nowhere to land — the density-ramp trap, now an operational reality rather than a scoping one. Hold too much and you are paying to idle megawatts. Forecast the ramp curve generation-by-generation, size the irreversible substrate to it, and keep the reversible IT fit-out matched to current draw — reserving the headroom you cannot retrofit (floor, water, ampacity) while deferring the spend you can. The interconnection and energization sequencing that feeds this forecast (50–100 MW tranches, energized capacity leading commissioned load) is in Chapter 3.2; the scoping-time version of the ramp decision is in Chapter 1.1.
Workload-aware operations and demand response
The most powerful lever an AI factory has over its own power profile is the scheduler. Because a large share of the fleet's load is interruptible (batch inference) or schedulable (RL rollouts, off-peak training), the operator can shape when and where power is drawn — not just react to it. Workload-aware operations is the practice of treating the power budget and the job queue as one optimization: placing interruptible work where stranded capacity hides, throttling low-SLO tenants first when a cap engages, smoothing the aggregate profile by phase-offsetting correlated jobs, and shifting deferrable load to cheap or low-carbon hours.
That same flexibility is what makes demand response and curtailment economically real for AI loads — and what makes it dangerous. The grid will pay handsomely for flexible load: Duke’s February 11, 2025 model estimates 98 GW of US load integratable at just 0.5% annual curtailment, with a 1–2% peak-demand cut translating into 0.5–2.8% rate reductions system-wide. For a facility, participating means accepting curtailment events (averaging ~2 hours) in exchange for a better interconnection deal or a demand-response payment — increasingly a condition of large-load tariffs, not an option. Everything then rides on which workloads you make eligible. Curtail batch inference or defer a checkpointable training run only after pricing deadline slack, lost work, restart, rebound demand and the contract’s baseline and notice requirements; checkpointability does not make the cost zero. Curtail an online-inference SLA or a synchronous run mid-step and you destroy goodput value that dwarfs the demand-response payment: a breached latency SLA, or a forced checkpoint-and-restart that throws away hours of correlated compute. The revenue is real, but eligibility must be contracted and tested for each workload. Online inference may participate through spare headroom or admission control only while preserving its SLO; synchronous training needs a validated checkpoint and restart budget.
| Workload | Interruption cost | DR eligibility | Mechanism on a curtailment event |
|---|---|---|---|
| Batch inference | Deadline delay, replay and rebound demand | Eligible only within priced deadline/recovery slack | Pause/defer the queue; resume after the event |
| RL rollouts | Low (staleness-tolerant, restartable) | Qualify trainer starvation, staleness and recovery cost | Throttle or pause rollout generation; protect the trainer |
| Synchronous training | High (forced checkpoint/restart) | Qualify checkpoint, restart and notice/deadline budget | Power-smooth/ramp-limit; checkpoint then idle if event is long |
| Online inference | SLO breach, credits and lost contribution if capacity is insufficient | Eligible with proved spare capacity, admission or geographic shift | Preserve latency/correctness SLO during reduced capacity and recovery |
Energy and efficiency operations
The final function folds power, thermal, and workload management into a single objective: useful work per megawatt-hour. PUE measures only the facility overhead — and at warm-water DLC's ~1.1 there is little overhead left to chase. The frontier has moved inside the white space, to metrics like work-per-watt and goodput-per-MWh that count whether the IT load itself is doing useful computation or burning power on idle GPUs, retries, throttled clocks, and failed jobs. A facility at PUE 1.1 running GPUs at 30% MFU wastes far more energy per unit of useful work than one at PUE 1.2 running at 45% — under matched hardware, peak-per-watt and workload, useful output per facility MWh follows productive-time fraction × conditional MFU ÷ PUE; multiply a separate active-duty factor only if Chapter 14.1’s clocks prove it is not already included, and the numerator moves over a far wider range than the denominator. Energy operations therefore reaches past the cooling plant into the cluster: raising MFU, cutting failure-induced rework, and timing flexible load to low-carbon, low-cost grid hours. By 2026 the efficiency levers that matter most are scheduling and reliability, not mechanical plant. → the full KPI and reliability-economics treatment is in Chapter 14.1.
Deep dive: the integrated operating loop — one budget, four control planes
The reason this chapter binds power, transients, stranded capacity, planning, and demand response into one subject is that in a live facility they are not separable — they are four control planes acting on a single shared resource, the energized power budget, and an action on one ripples through all four. Make this concrete with a single event: a 50 MW training run finishes and frees its budget at 2 a.m.
Capacity planning sees freed headroom and the placement engine looks for work to fill it. Workload-aware ops routes a deferred batch-inference queue into the gap — but that queue is being held at low power precisely because a demand-response event is forecast for the morning peak, so the scheduler front-loads it overnight to be done before curtailment. As the batch load ramps, the transient-management layer applies ramp-rate limiting so the 50 MW does not step back in one swing, and the power-cap hierarchy holds the row budget while the freed training racks and the new batch racks briefly overlap. Thermal ops tracks the heat load shifting from one row to another and rebalances flow so neither row drifts toward its throttle threshold. Stranded-capacity diagnostics note that the batch work filled power headroom but left cooling headroom unused in an adjacent row — a fragmentation signal for the next placement pass.
No single dashboard owns that sequence; it is the integrated operating loop running continuously. The organizational consequence — that facility-ops and ML-platform-ops must share one view of the budget and one definition of goodput, rather than optimizing their halves in conflict — is the operating-model question taken up in Chapter 14.11, and the change-control discipline that governs any deliberate move within this loop (a cap change, a setpoint change, a DR enrollment) is the MOP/SOP regime in Chapter 14.12.
Cite this chapter
Fehn, J. (2026). Capacity, Power & Thermal Management in Operation (Chapter 14.7). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-7-capacity-power-and-thermal-management-in-operation (accessed 2026-09-29).
@misc{aidc-14-7,
author = {Fehn, Jacob},
title = {Capacity, Power & Thermal Management in Operation (Chapter 14.7)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-7-capacity-power-and-thermal-management-in-operation},
note = {Accessed 2026-09-29}
}