The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Sustainability & Efficiency › 15.2

Chapter 15.2

In this chapter · 8 sections
Term help

Energy Efficiency: Cooling, Free Cooling, Setpoints & Power-Chain Losses

Coolant temperature, economizer type, and voltage class are irreversible efficiency decisions — bank them at scoping or pay in megawatts every hour; the leverage now sits in setpoints and the power chain.

POWER-BOUNDGOODPUT

What you'll decide here

  1. The facility-water and coolant setpoint band (chilled ~18-27 C vs warm-water ~32-45 C, i.e. the ASHRAE W17/W27 versus W32–W45 facility-water classes, not the identically numbered air-inlet range) — a decision that, together with climate, approach temperatures, rejection method and part-load curves, determines compressor-free hours and the heat-pump duty of reuse.
  2. Whether to design for free cooling / economization at all, and which kind (airside, waterside, or dry-cooler-direct) — a climate-and-water-coupled fork that sets your annualized PUE floor, not the nameplate.
  3. The IT inlet and coolant-return setpoints you will actually operate at — unused headroom can cost compressor energy where the selected plant curve permits a higher setpoint, but raising setpoints narrows the thermal ride-through margin that protects 1 kW+ GPUs.
  4. The power-chain topology and UPS operating mode (double-conversion ~94-97% vs bypass-ECO or active high-efficiency ~98-99%; 415/480 VAC vs an 800 VDC path) — 2-5 points of end-to-end efficiency that compounds across every watt, for the life of the building.
  5. Whether ML-driven cooling control and part-load-aware operation are in the design basis or retrofitted later — the difference between a plant that is efficient at nameplate and one that is efficient at the measured part-load bins where it actually operates.

The efficiency conversation in AI data centers is usually conducted in the wrong units. People argue about PUE as if it were a scoreboard, when it is a consequence: the visible residue of a dozen upstream decisions about coolant temperature, economizer design, setpoint discipline, and how many conversion stages sit between the utility and the GPU's voltage regulator. Most of the leverage is spent at design time, in decisions that are expensive or impossible to reverse once steel is cut and water is plumbed.

Three things changed with the AI buildout. First, the denominator grew: an air hall's cooling overhead used to be 40-50% of IT load; a well-designed warm-water direct-to-chip (DLC) plant can push facility overhead toward 5-10%, which means the marginal efficiency game has moved off the chiller and onto the coolant loop and the power chain. Second, density removed the slack: with qualified designs spanning Supermicro's 60 kW air-cooled HGX B300 layout, HPE's 132 kW nominal GB200 NVL72, NVIDIA's up-to-142 kW GB300 facility design point, and NVIDIA's 330 kW Vera Rubin cabinet design basis, you no longer have the option to waste a few points of efficiency on conservative setpoints — the cooling plant is on the critical path, and every avoidable watt of overhead is a watt of IT capacity you cannot energize in a power-bound site. Third, the metric stopped meaning what it used to: PUE quietly rewards moving losses inside the IT envelope — the boundary critique of Chapter 15.1. That chapter owns the metric stack; this one treats the physical decisions that move it.

The efficiency budget: where the leverage actually is

Start by drawing the energy budget honestly, because that is where you find the leverage. In a modern AI facility, non-IT (overhead) energy splits roughly into cooling (60-80% of overhead) and the power chain (conversion and distribution losses, most of the rest), with lighting and ancillaries a rounding error. Cooling is the dominant lever, and within cooling the dominant sub-lever is how many hours per year you can avoid running the compressor at all, rather than the chiller's efficiency at full load. The highest-value efficiency decision is therefore not a better chiller but a coolant loop warm enough that the chiller is mostly off.

The thermodynamics explain the leverage. Mechanical (compressorized) cooling is the expensive mode; free cooling — rejecting heat directly to ambient air or water without a vapor-compression cycle — is nearly free by comparison. The fraction of the year you spend in each mode is set almost entirely by two numbers you choose at design time: the coolant/facility-water supply temperature and the approach temperature of your heat-rejection equipment relative to ambient. Raise the supply temperature and you convert thousands of compressor-hours into economizer-hours. It is why a Reykjavík facility and a Phoenix facility built to the same nameplate PUE can have wildly different annualized PUE. The nameplate is a snapshot at design conditions; the annualized number is the integral over the weather, and the weather is set by siting (Chapter 3.7) and exploited by setpoint.

Cooling architecture and the energy paths it selects

Cooling architecture is the first fork: it selects the heat-rejection path and compatible rack densities before you tune a single setpoint; climate, loading and plant curves then determine annual PUE. The choice comes in a small set of discrete regimes rather than a smooth dial, each with a different parasitic load and a density fit set by the named OEM configuration (the cooling cliff of Chapter 5.1). The table below names the physical mechanisms and equipment limits; compare their annual energy at the same climate, load and useful service before choosing. Cooling architecture alone does not supply a portable PUE ranking.

Cooling architecture → achievable PUE band and the efficiency mechanism
Cooling architectureDensity fitPUE bandFree-cooling exposureWhy it sits there
Legacy air + DX/chiller≤ ~30 kW/rack~1.4–1.6Low — compressor-bound much of the yearCompressorized cooling dominates; warm supply air limited by IT inlet specs
Optimized air + economizerup to 60 kW/rack in a named OEM design (Supermicro HGX B300)~1.08–1.15 in hyperscale free-cooling fleets (Meta 1.08, Google 1.10, Microsoft/AWS ~1.15); higher where economizer hours are fewerModerate — airside/waterside economizer hoursContainment + economization cut compressor hours; still air-transport-limited
Rear-door HX / air-assisted liquid~40–75 kW/rack~1.2 (CIBSE hybrid liquid/air example)Moderate-high — warmer loop enables more free hoursLiquid-to-the-door removes the air-transport penalty; loop can run warmer
Direct-to-chip liquid (warm-water)~75–200+ kW/rack1.05–1.15 (design illustration)High — 32–45 °C loop is dry-cooler / economizer friendlyHeat captured at the source in a warm loop; compressor often unnecessary
Single-/two-phase immersion~50–200+ kW/rack1.01–1.03 (GRC single- vs two-phase comparison)Very high — high-grade warm bath rejects to ambientMinimal fan/pump parasitics; best floor, but fluid/PFAS and serviceability costs
PUE bands are cited 2025 design ranges — screening, not measured: Uptime’s DX example and the ~1.6 industry average (legacy air), SemiAnalysis’s February 2025 hyperscaler fleet figures, CIBSE’s October 2025 hybrid example, SemiAnalysis’s warm-water DLC illustration and GRC’s immersion brochure; no common IT-energy boundary, load curve or measured population sits behind them. The optimized-air density point is Supermicro’s four-node, 60 kW air-cooled HGX B300 layout. Annualized PUE lands inside a band as a function of climate and setpoint discipline — the colder the site and the warmer the loop, the closer to the floor — but expose weather bins, load, heat-capture fraction, power-chain losses and meter locations before contracting against a number; a zero-evaporation dry-cooled DLC plant in a hot basin annualizes above its temperate band unless adiabatic assist trims the peak hours, the water-for-PUE trade priced in Chapter 15.4.

The architecture you can choose is constrained by rack density and vendor qualification on the left; the parasitic-energy terms are on the right; and free-cooling exposure is the lever the rest of this chapter pulls. The consequence is sharp: choosing air near its qualified limit instead of warm-water DLC preserves higher fan power and narrows the economizer window; climate and setpoints determine the annualized PUE cost. The cooling architecture is, in effect, a decision about how many compressor-hours you have signed up to pay for over the building's life.

Free cooling and economization: the highest-value efficiency decision

Free cooling — economization — is the act of rejecting heat to the environment without running a vapor-compression cycle. It is the single largest annualized-efficiency lever in the building, and it comes in three architectures that trade water, capital, and reachable hours against one another. The fork is whether reachable compressor-free hours repay the economizer’s installed and maintenance cost, then which kind, because each one couples to a different siting constraint and a different downstream cost.

Airside economization pulls filtered outside air directly into the hall when ambient is cool and dry enough, exhausting hot air rather than recirculating and re-cooling it. It is the cheapest to operate and the most water-free, but it imports outdoor humidity, particulates, and gaseous contaminants into the white space — a real reliability liability — and it works only when the air-cooled IT inlet spec (ASHRAE A1-A4) can be met directly. Waterside economization keeps the air loop sealed and instead uses a cooling tower or dry cooler to chill the facility water without the chiller, via a plate heat exchanger when wet-bulb (or dry-bulb) is low enough. It tolerates a wider climate envelope and keeps contaminants out, at the cost of water (evaporative towers) or a larger dry-cooler footprint and a warmer achievable loop. Dry-cooler-direct (compressor-less) operation is the warm-water DLC endgame: if the selected facility loop is validated for dry-cooler operation, a dry cooler can reject to ambient air across most of the year with no evaporation and no compressor at all — zero process water for cooling, and PUE that approaches the pump-and-fan parasitic floor.

Free-cooling architecture → the trade it forces
Economizer typeSealed white space?Water useReachable free hoursPrimary downside
Airside (direct outside air)No — outside air enters hallNone (unless adiabatic assist)High in cool/dry climatesImports humidity, particulates, gaseous contaminants; needs filtration + RH control
Airside + adiabatic assistNoModerate (evaporative pre-cool)Extends warm-climate hoursReintroduces water; adds spray/media maintenance
Waterside (tower + plate HX)YesHigh (evaporative) or low (dry)High where wet-bulb is lowTower water, blowdown, Legionella control (ASHRAE 188)
Dry-cooler-direct (warm loop)Yes≈ Zero for coolingVery high if loop ≥ ~40 CLarger heat-rejection footprint; needs warm-water DLC to begin with
Reachable-hours depends on climate; figures are directional for a temperate-to-cool site. Engineering home for heat-rejection equipment is Chapter 5.8; the warm-water facility loop is Chapter 5.7.

The trade is a water-versus-PUE-versus-capital triangle, and the AI-density era has bent it decisively toward the dry-cooler-direct corner. The reason is the warm loop: once DLC lets you capture heat at a selected warm-water setpoint instead of cooling air to 18-27 C, the dry cooler becomes viable for most of the year in most temperate climates, which lets you design water out of the building entirely for cooling through dry, non-evaporative heat rejection, the near-zero-WUE condition hyperscalers now publish (Microsoft’s December 9, 2024 design announcement estimates avoiding more than 125 million liters per datacenter per year; it scheduled Phoenix and Mt. Pleasant pilots for 2026 and first new-site operation for late 2027). The same decision governs water stewardship in Chapter 15.4, which is why efficiency and water cannot be optimized separately: the coolant setpoint that maximizes free-cooling hours is also the one that lets you eliminate evaporative water. The cost you pay is footprint and capital — dry coolers are larger and cannot reach as low a loop temperature as an evaporative tower on a hot day — which is precisely why this is a siting-coupled decision, not a mechanical one.

Warm-water cooling: the facility-water setpoint

Everything above converges on one number: the facility-water supply temperature. The industry has spent decades over-cooling — treating the 18-27 C recommended air-inlet range as if it prescribed chilled-loop temperature, although air inlet and facility water are different measurement boundaries — and in doing so threw away both free-cooling hours and any chance of heat reuse. Warm-water DLC inverts the logic. ASHRAE's liquid-cooling classes (W17, W27, W32, W40, W45, and W+ in the 5th-edition Thermal Guidelines) are keyed to the upper facility-water supply temperature precisely to make this a deliberate design choice, while the OCP October 1, 2024 coolant roadmap proposes 30°C at the silicon/TCS interface to limit expensive future cooling retrofits. That design target is separate from NVIDIA’s Vera Rubin reference: 41°C FWS into the CDU and 45°C rack coolant. Neither is a universal facility-loop setpoint.

The consequence chain from this one setpoint is the most important in the chapter. A warmer loop (a) widens the temperature difference between your coolant and ambient, which (b) lets a dry cooler or economizer reject heat across more of the year, which (c) can drive compressor-hours toward zero where weather and plant curves permit, which (d) drops annualized PUE toward the parasitic floor, and simultaneously (e) lifts the return-water temperature high enough that the waste heat becomes a sellable product instead of a disposal problem (Chapter 15.5). One supported setpoint, five downstream effects to test: fan and pump demand can offset compressor savings, and a hotter return earns money only when the heat buyer accepts and pays for it. The cost is margin: a warmer loop leaves less thermal headroom, so the cold-plate design, flow rate, and CDU approach temperature must be tighter and the controls more disciplined (the transient-stability problem of Chapter 5.12). The engineering of the loop that delivers it lives in Chapter 5.7.

Setpoint strategy: every conservative degree is a watt you chose to spend

Setpoint strategy is where design intent meets operating reality, and it is where most facilities quietly leave efficiency on the table. The instinct of an operations team is to run cold and conservative — it feels safe. But ASHRAE's recommended IT inlet envelope has been 18-27 C for years, with A1-A4 allowable ranges extending to ~32-45 C, and every unused degree can cost compressor energy, while warmer inlet air can raise server-fan power; spend the thermal margin only after checking total facility energy and accepted work. The setpoint decision is therefore a deliberate risk-versus-efficiency trade, and it must be made explicitly rather than defaulted to "cold."

The fork has three settings. Conservative (low IT inlet, cold loop, wide margin) maximizes ride-through and equipment longevity headroom at the cost of free-cooling hours — defensible for an air hall whose equipment limits or poor thermal monitoring require that reserve, and where the tested fault response cannot support a warmer inlet. ASHRAE-recommended (mid-band inlet, moderate loop) is the safe default for most operators. Aggressive / allowable-band (high inlet within A-class limits, warm loop) maximizes economizer hours and heat-reuse grade, and is a candidate for a well-instrumented liquid-cooled facility whose OEM envelope and cooling-loss tests support it — but it narrows the thermal ride-through window, which matters enormously when a 1 kW+ GPU's response to a cooling-loss event is product- and implementation-specific and must be established by OEM data or tests (the resilience coupling of Chapter 12.2). The downstream cost of running warm is not normally efficiency — it is that you have spent your transient margin, so your pump redundancy, UPS-backed cooling, and controls stability had better be commissioned to match. Raising setpoints without first proving ride-through is how an efficiency initiative becomes an outage.

Power-chain efficiency: the other half of overhead

Cooling is the larger lever, but the power chain compounds across every watt the building draws. The four-stage AC fixture in Chapter 4.1 gives 80.4–86.7% utility-to-VRM, with a representative 82.2%; that chapter owns its loss ledger and the additional distribution-loss boundary. The separately sourced SemiAnalysis SST/800 VDC model (~87%, May 26, 2026) is not a guaranteed improvement on that guide fixture. A like-for-like difference requires the same load point, included stages and final output boundary.

Two power-chain decisions carry most of the leverage. The first is the UPS operating mode: a double-conversion UPS runs at ~94-97% efficiency because it rectifies and re-inverts continuously; high-efficiency modes reach ~98-99%, but they are two different products. Bypass-style ECO parks the inverter on standby and passes utility power through, so a transfer back to inverter costs a few milliseconds that must be reconciled against the load's ride-through requirement. Active high-efficiency modes keep the inverter energized, hold conditioning and reactive-power compensation, and transfer to double conversion without a break. Contemporary double-conversion designs also reach ~97%, so size the recoverable gain from the two candidates' measured load-dependent efficiency curves and acceptance tests, not from the mode's name (the UPS architecture decision lives in Chapter 4.5). The second is the voltage architecture: the 48 V → ±400 V → 800 VDC transition (Chapter 4.7) removes conversion stages, lets the same wire gauge carry ~157% more power at 800 VDC than at 415 VAC in NVIDIA's three-wire-DC versus four-conductor-AC comparison (Oct 2025), and is the structural enabler for 600 kW+ racks — which is why the efficiency case and the density-ramp case point the same way. Both decisions are substantially irreversible: you commission a UPS topology and a voltage class once, and re-doing either mid-life is a rebuild, not a tune-up.

1.54
Uptime Institute 2025 survey-respondent average annualized PUE (n=681); not an industry-weighted fleet statistic
Scope & caveats

Respondent-survey statistic, not an industry-weighted fleet average. The public claim does not provide a population-weighting method; use it as a dated survey benchmark and retain the sample and survey methodology when comparing it with a named facility or fleet.

2025 survey figure (n=681). Uptime's 2026 survey (published 2026-07-28; analysed 2026-08-06) reports two separately defined successors: 1.52 as the annual survey average and 1.36 on a capacity-weighted basis — the gap is the larger-facility advantage, not an improvement in the same population. Neither is a measured average of the world fleet. Same 2026 survey: modal installed-base rack density reached 11 kW (from 9 kW in 2025) — the installed base, not the AI-factory design point.

1.09
Google Q2 2026: fleet TTM PUE 1.09 versus quarterly 1.10; best individual-site TTM 1.04 at Central Ohio (Lancaster)
Scope & caveats

Reporting period ended June 30, 2026. Fleet TTM, fleet quarter and individual-site TTM are separate statistics, not a cooling-architecture trial.

~1.4–1.6estimate
legacy air-cooled PUE reference band: Uptime DX example ~1.40 to the ~1.6 industry average
Scope & caveats

Cited screening band, not a measured population; a site’s legacy-air result follows its compressor hours and containment.

1.08–1.15
hyperscaler economized-air fleet PUE as quoted (Meta 1.08, Google 1.10, Microsoft/AWS ~1.15)
Scope & caveats

Company-reported fleet averages as quoted; economized air-cooled halls across mixed climates, not a design guarantee for one site.

1.05-1.15estimate
Warm-water DLC design illustration, not a measured architecture-wide range; use matched annual climate/load and meter boundaries
Scope & caveats

No matched measured population establishes an architecture-wide 1.05–1.15 envelope. This cited design illustration is not an annual guarantee; climate/load bins, rack heat capture, power-chain losses and meter boundaries govern a project’s result.

~60-80%
cooling share of non-IT (overhead) energy — the dominant efficiency lever
~41 °C FWS -> 45 °C rack coolant
NVIDIA’s named 41°C FWS/45°C rack-coolant example can support dry rejection when the local approach and ambient envelope permit it; OCP’s October 2024 roadmap separately proposes 30°C at the silicon/TCS interface
Scope & caveats

Published NVIDIA reference design point for the Vera Rubin POD, a shipping platform: a named FWS/CDU/TCS example, not a universal facility-loop standard and not a metered facility loop. Derive the site setpoint from the named rack's approved envelope and the CDU approach.

94-97% → 98-99%
UPS efficiency: double-conversion vs bypass-ECO / active high-efficiency mode (recoverable points are product- and load-dependent)
Scope & caveats

Bypass-style ECO (inverter on standby, short break on transfer) and active high-efficiency modes (inverter energized, conditioning held, uninterrupted transfer) are different products that both reach ~98-99%. Contemporary double-conversion designs reach ~97%, so the recoverable delta is product- and load-dependent — take it from measured efficiency curves and acceptance tests, not from the mode name.

45 °C maximum liquid inlet; 65 °C maximum liquid return (separate limits)
QCT GB200 NVL72 QoolRack reference maxima: 45 °C liquid inlet and 65 °C liquid return; separate limits, not a selected operating pair
Scope & caveats

Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.

Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.

ML-driven cooling optimization: closing the part-load gap

A plant designed to be efficient at nameplate is not the same as a plant that operates efficiently, because the building almost never sits at nameplate. AI facilities can spend long periods below nameplate, with training jobs that ramp and checkpoint and inference fleets that swing with arrivals; use the measured load bins to size the annual penalty — and cooling plant that is efficient at full load is frequently inefficient at part load, where pumps, fans, and chillers run off their best-efficiency point. The part-load gap is real money, and it is the domain where machine-learning control has earned its place in the design basis rather than as an afterthought.

The canonical result is Google's July 2016 announcement, where DeepMind machine-learning recommendations for cooling-plant setpoints cut cooling energy by ~40% and PUE overhead by ~15% against a human-tuned baseline — an advisory system, not the autonomous closed-loop control Google described separately in 2018 with its own results. Read it as historical proof of the mechanism against a 2016 air-cooled plant, not as the savings still on the table in an already-optimized 2026 liquid plant — the kind of gain that is invisible to nameplate PUE because it lives entirely in the part-load, multi-variable interactions a static setpoint table cannot capture. The mechanism is straightforward: a cooling plant has dozens of coupled actuators (pump speeds, valve positions, tower fan speeds, chiller staging) and a non-linear response surface that shifts with load and weather; an ML controller searches that surface continuously where a human operator sets-and-forgets. The decision is whether to design for this — instrumenting the plant densely enough (the DCIM telemetry of Chapter 14.2) and giving the controller safe authority — or to bolt it on later against a plant that lacks the sensors and actuators to exploit it. Retrofitting observability is far more expensive than designing it in, which is why agentic and RL-based control (Chapter 14.13) belongs in the efficiency design basis, not the wish list. The caution: an ML controller that optimizes facility PUE without a goodput constraint optimizes the wrong thing in the wrong direction. With facility losses roughly fixed, cutting IT load raises PUE, so the objective rewards keeping servers drawing power, and idle-but-powered IT flatters the ratio while producing nothing. The objective must be useful-work-per-watt, with the GPU thermal envelope as a hard constraint, never PUE in isolation.

Deep dive: the four delta-Ts and why each one is an efficiency decision

Practitioners decompose a liquid-cooled facility's thermal path into a chain of temperature differences — the "four delta-Ts" — and each one is a lever you trade against capital, parasitic power, and free-cooling hours. Walking them from chip to sky makes the efficiency physics concrete.

Delta-T #1: chip-to-coolant (across the cold plate). Set by cold-plate thermal resistance (~0.02-0.03 C/W) and flow rate. A tighter delta-T here lets the coolant run warmer for the same junction temperature — but demands more flow, which costs pump power. Delta-T #2: coolant-loop rise (the ~7.5-12 C the coolant gains across the rack). A larger rise means less flow for the same heat — lower pump parasitics — but a higher return temperature that the CDU must handle. Delta-T #3: CDU approach (the ~3-5 C the heat exchanger loses transferring from the technology-cooling loop to the facility-water loop). Smaller approach means a warmer facility loop for the same chip temperature — directly more free-cooling hours — but a larger, costlier heat exchanger. Delta-T #4: facility-loop-to-ambient (the margin the dry cooler or tower needs over wet- or dry-bulb to reject heat). This is the one the weather controls and the warm loop widens; it is the delta-T that decides how many hours the compressor stays off.

The unifying insight: efficiency is the art of spending the right delta-T in the right place. Every degree you can push into the facility-loop-to-ambient delta-T (#4) by tightening the upstream three (#1-#3) is a degree of free-cooling exposure. Warm-water design, cold-plate engineering, and CDU sizing therefore form one continuous budget, and the budget's bottom line is annualized PUE. Engineering homes: cold plates and flow in Chapter 5.4, the CDU and secondary loop in Chapter 5.6, heat rejection in Chapter 5.8.

The PUE caveat: raise setpoints against TUE or work-based numbers, not facility PUE

PUE's boundary blind spot — server fans and pumps count as useful IT the moment they sit inside the IT envelope, so direct-to-chip liquid can make a genuinely more efficient facility look worse on PUE than the air hall it replaced — is derived in full, with the TUE = ITUE × PUE fix, in Chapter 15.1. The operational instruction for this chapter: when you raise setpoints or warm the loop, watch a TUE-class or work-based number, not facility PUE — otherwise you can congratulate yourself on a falling PUE while server fans spin up and goodput falls.

Part-load and utilization: efficient at the load you actually run

Annual controls comparison — sum each bin before dividing
Bin / durationIT MWOverhead MW: baseline / proposedFacility MWh: baseline / proposed
Cool / 4,000 h80.8 / 0.635,200 / 34,400
Mild / 3,000 h101.2 / 0.933,600 / 32,700
Hot / 1,760 h102.0 / 1.621,120 / 20,416

Unsupported inputs are stated in the preceding callout. Each cell’s energy is duration × (IT + overhead); these are synthetic bins, not a weather file.

Trace and action. Annual IT energy = 4,000 × 8 + 3,000 × 10 + 1,760 × 10 = 79,600 MWh. Sum the table: baseline facility energy is 89,920 MWh, proposed 87,516 MWh. Annual PUE is respectively 89,920/79,600 ≈ 1.13 and 87,516/79,600 ≈ 1.10, with about 2,400 MWh/year saved. Carry the proposed mode into the operating budget only after useful-work and failure-state acceptance; an acquisition decision still needs its delivered-price and cost inputs below.

Flip. Cool and mild bins save 800 + 900 = 1,700 MWh. If proposed hot overhead is h MW, annual saving = 1,700 + 1,760 × (2.0 − h). Equality is h = 2.0 + 1,700/1,760 ≈ 3.0 MW. At the stressed 3.2 MW, proposed annual energy becomes 90,332 MWh, about 400 MWh more than baseline: reject that mode. The actual quality, thermal or water limit can reject it sooner. Chapter 15.1 owns the annual ratio method; Chapter 5.8 supplies the real climate and plant curves, and Chapter 13.5 proves the proposed controls.

The last efficiency lever is the most often ignored because it does not appear on a nameplate at all: match the plant's efficiency curve to the load profile you will actually run, not the design-day peak. A cooling plant or UPS that is most efficient at 100% load is the wrong plant for a facility whose measured load trace spends most hours below peak. The fix is modularity and staging — multiple smaller chillers, pumps, and CDUs that can be staged on and off to keep the running units near their best-efficiency point, rather than a few large units running inefficiently part-loaded. The same logic governs the power chain: a UPS bank sized so that each module sits in its high-efficiency band at typical load beats one sized so every module idles at 30%.

This is a design-time decision with an operating-time payoff, and it interacts with two threads of this guide. On the power-bound axis, part-load-efficient plant means more of your contracted megawatts reach the IT load instead of being lost to oversized, under-loaded equipment — directly more compute per interconnection slot. On the goodput axis, the staging logic must never compromise the thermal ride-through that protects the GPUs, so part-load efficiency and resilience are co-designed (Chapter 14.7 on operational capacity/thermal management). The anti-pattern is the facility designed and benchmarked at peak, commissioned with a single efficiency point, and then operated for years at a load where that point is irrelevant — efficient on paper, wasteful in production.

Release the energy investment on the annual bins, then prove the setpoint. Carry the roughly 800 MWh/year saving from Chapter 15.1, and the 2,400 MWh/year computed above, only as illustrative energy results. For this project, obtain each bin’s metered energy and delivered tariff from Chapter 3.3; subtract added maintenance, water treatment, controls and rollout costs, then price the resulting dated cash flows against installed capex using Chapter 1.8. A saving that disappears when demand charges or warm-day IT fan power enter the boundary rejects the investment.

The mechanical lead proposes the FWS and TCS supply, ΔT, flow, fluid and CDU approach for the named rack. The commissioning lead must acquire traces for design-day operation, a pump or CDU loss, sensor failure and controller fallback, with junction temperature, IT/facility power and accepted-output limits stated before testing. Until those traces exist, hold the warmer setpoint. The efficient normal point is no bargain if a cooling fault consumes the remaining ride-through time before the controller can shed load; Chapter 13.5 owns the acceptance test.

This chapter is the strategy-and-decisions layer over the cooling and power engineering treated in depth elsewhere. The metric stack it leans on — PUE, WUE, ERF, TUE, work-based metrics — is built in Chapter 15.1. The density wall that gates cooling architecture is Chapter 5.1; warm-water facility loops in Chapter 5.7; heat rejection, economizers, dry coolers and towers in Chapter 5.8; the CDU and secondary loop in Chapter 5.6; DLC cold-plate engineering in Chapter 5.4; and setpoint/controls transient stability in Chapter 5.12. The power-chain decisions live in Chapter 4.5 (UPS modes) and Chapter 4.7 (the DC power revolution). Climate-and-water siting is Chapter 3.7; the efficiency-vs-water coupling is Chapter 15.4; the warm-loop heat-reuse payoff is Chapter 15.5. The goodput-vs-availability reframe that disciplines setpoint risk is Chapter 12.2; operational capacity/thermal management is Chapter 14.7; DCIM telemetry is Chapter 14.2; and ML/agentic cooling control is Chapter 14.13.
Cite this chapter
Fehn, J. (2026). Energy Efficiency: Cooling, Free Cooling, Setpoints & Power-Chain Losses (Chapter 15.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-15-sustainability-and-efficiency/15-2-energy-efficiency-cooling-free-cooling-setpoints-and-power-chain-losses (accessed 2026-09-29).
@misc{aidc-15-2,
  author       = {Fehn, Jacob},
  title        = {Energy Efficiency: Cooling, Free Cooling, Setpoints & Power-Chain Losses (Chapter 15.2)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-15-sustainability-and-efficiency/15-2-energy-efficiency-cooling-free-cooling-setpoints-and-power-chain-losses},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit