Chapter 0.3
In this chapter · 6 sections
Vocabulary, Mental Models & the Metric Stack
Vocabulary collisions — power figures, capacity units, metrics, nines, network tiers — cost real money downstream; fix a shared reading of each term before any contract or sizing decision depends on it.
What you'll decide here
- Which power number anchors every contract and capacity claim — critical IT power, total facility power, or grid draw — because conflating them mis-sizes the interconnect, the switchgear, and the lease by the coincident non-IT load and meter boundary.
- Which capacity unit you scope, budget, and order against — the rack, the NVLink/scale-up domain, the scalable unit (SU), the pod, or the gigawatt campus — and the blast radius each implies.
- Which single number the facility is being optimized for: facility efficiency (PUE/WUE), useful work (MFU/goodput), or unit economics ($/GPU-hr, $/M-tokens) — they pull in different directions.
- How you will write and read availability claims — the nines, their downtime budgets, and serial-vs-parallel composition — so a vendor's '99.99%' cannot quietly mean four different things.
- Which of the three network tiers (scale-up, scale-out, scale-across) a given bandwidth, latency, or oversubscription claim refers to — because the same word means radically different physics at each tier.
Two parties using the same word for different quantities is the recurring failure mode in AI-data-center projects, and the gap usually surfaces only after a contract is signed or a slab is poured. A developer quotes "100 MW" and means grid draw; the tenant hears critical IT power and sizes a cluster against capacity that the building overhead also needs. A neocloud advertises "99.99% uptime" and means a single node's hardware availability; the customer assumed it described the training job's goodput. A vendor cites "1.8 TB/s" of bandwidth without saying it is the scale-up domain, and a network architect budgets the scale-out fabric against it. None of these are lies. They are vocabulary collisions, and each one costs real money downstream.
This chapter fixes the shared language. We pin down the power vocabulary (the three power numbers that are routinely conflated), the capacity units (from the rack up to the gigawatt campus), the metric stack (efficiency, useful-work, and unit-economics families and how they trade off), the availability algebra (the nines, their downtime budgets, and how components compose), and the three-tier network hierarchy (scale-up vs scale-out vs scale-across). This is an orientation chapter, not the canonical home for any of these — each metric's rigorous definition, measurement plan, and gotchas live in their discipline chapters, and we point you there. The job here is to make sure that when Part 3 says "facility power" or Part 8 says "oversubscription" or Part 12 says "goodput," you already read the word the same way the author wrote it.
Power vocabulary: the three numbers everyone confuses
Start with power, because power is the binding constraint of the 2026 era and the most frequently mis-stated number in the building. There are three distinct power figures along the chain, and a project that does not keep them separate will mis-size something expensive.
Critical IT power is the power delivered to the compute itself — the racks, after the last PDU, the number the workload actually consumes. Total facility power adds everything required to keep that IT alive: cooling plant, pumps, fans, lighting, losses in the UPS and transformers, the whole non-IT overhead. The ratio between them is PUE, which ISO/IEC 30134-2 defines as total facility energy over IT energy across the same continuous 12-month period; the power-ratio form everyone quotes is a design or interim derivative, not the annual metric. Grid draw (or utility demand) is net import at the meter, and contracted capacity is the MW limit the interconnection agreement sets — a limit, not a reading, and on-site generation or storage separates both from facility load. Take a stated peak-load case with 100 MW critical IT plus 20 MW of coincident non-IT overhead: that requires 120 MW of facility power. Its 1.2 instantaneous ratio is not proof that annual PUE 1.2 supplies the same peak capacity. Size the service from the peak operating cases, and study inrush and transients separately rather than folding them into an unexplained headroom allowance. Quote the wrong one of these in a lease or a power-purchase term sheet and the error propagates into switchgear ratings, transformer orders, and the size of the queue slot you reserve.
Two further distinctions trip people up. MW vs MVA vs kVA: real power (MW) is what does work and what you pay for in energy; apparent power (MVA/kVA) is the magnitude of complex power, with S² = P² + Q² for sinusoidal steady state, and transformers and generators commonly carry kVA/MVA ratings, while switchgear is normally specified by voltage, continuous current, interrupting/withstand ratings and the applicable standard. At a power factor of 0.9, a 100 MW load is ~111 MVA of apparent power — so equipment sized in MW to match the load will be undersized by the inverse of the power factor. AI loads, with their large rectifier front ends and sharp synchronous transients, make power factor and the MW/MVA gap a live design concern. → Chapter 4.1. And TDP vs electrical design power (EDPp): a GPU's thermal design power (TDP) is the per-chip steady-state envelope (~700 W for H100, ~1.0-1.2 kW for a Blackwell GPU), while the figure that actually sizes your power chain is the per-rack and per-cluster draw under real workloads, including transients that can swing tens of percent on synchronous all-reduce boundaries. Design to nameplate TDP summed naively and you will both over-provision average capacity and under-provision for the transient. → Chapter 4.5.
Capacity units: from the rack to the gigawatt campus
The second piece of shared vocabulary is the unit of capacity you scope, budget, order, and reason about failures in. AI facilities are built and discussed at five nested scales, and choosing the wrong granularity for a given decision is itself a mistake — you procure in scalable units, you reason about blast radius in fault domains, and the two do not always align.
The rack is the atomic power-and-cooling unit: an HPE GB200 NVL72 rack has a 132 kW nominal rack TDP, weighs ~1.47 t wet (3,245 lb fully loaded), and is the thing the slab, the busway, and the cooling manifold are sized against. Do that sizing against the generation you will actually deploy. GB300 NVL72 shipped in 2025 and is deploying through 2026 alongside GB200; Lenovo rates it at 135 kW rack TDP and up to 155 kW peak, while NVIDIA's full-rack facility design basis is up to 142 kW. The VR200 NVL72 entered full production in August 2026; Pegatron rates it at 188 kW Max Q / 228 kW Max P, while NVIDIA DSX sets a 330 kW cabinet facility design basis. The NVLink / scale-up domain is the set of accelerators with load/store addressability over NVLink — 8 in an HGX node, 72 in an NVL72, heading to 144 per rack in the announced Kyber generation — and to 576 across the eight-rack, optically-stitched Rubin Ultra NVL576 system of 72-GPU MGX racks (a different enlargement from the single ~600 kW Kyber NVL144 rack, whose eight-rack system is NVL1152). The GB200 NVL72 has 13.4 TB of GPU memory across its 72-GPU domain; the GB300 NVL72 raises that to 20 TB. This is the most important unit for the software, because tensor- and expert-parallelism must fit inside it to run at full bandwidth. The scalable unit (SU) is the repeatable building block of the cluster — a fixed bundle of racks, leaf switches, cooling, and power that the design replicates N times to scale; it is the unit you actually procure and commission, and the unit a reference architecture (DGX SuperPOD and its peers) is specified in. The pod is a deployment block of multiple SUs sharing a spine layer and often a fault domain. The gigawatt campus is the 2026 frontier unit of strategic conversation — multi-building sites in the 1 GW-plus class that are now the headline of hyperscaler announcements and the scale at which power, water, and grid impact become regional questions. → reference architectures in Chapter 0.4; the SU as a budgeting artifact in Chapter 0.2; fault domains in Chapter 12.1.
| Unit | Scale (2026 reference) | Bound by | Reason about it for |
|---|---|---|---|
| Rack | 1 rack; GB300 NVL72: up to 142 kW facility design basis; VR200 NVL72: 330 kW facility design basis; Rubin Ultra / Kyber: ~600 kW roadmap planning point (H2 2027) | Slab loading, busway, cooling manifold, one CDU's reach | Power & cooling density; floor loading |
| NVLink / scale-up domain | 8 (HGX) → 72 (NVL72) → 576 (Rubin Ultra NVL576) GPUs; Kyber NVL144: 144/rack, 1,152 per 8-rack NVL1152; each NVLink domain is load/store-addressable (13.4 TB GB200 / 20 TB GB300) | Copper/NVLink reach; switch radix | Whether tensor-/expert-parallelism fits at full bandwidth |
| Scalable unit (SU) | Repeatable bundle: racks + leaf switches + power + cooling | The reference architecture's replication block | Procurement, commissioning, the capacity ramp |
| Pod | Multiple SUs sharing a spine layer | Spine radix; a shared fault domain | Scale-out topology and blast radius |
| Gigawatt campus | Multi-building site, 1 GW-plus critical load | Grid interconnection; regional water/power | Siting strategy, grid impact, financing |
The metric stack at a glance
There is no single "efficiency" number for an AI factory, and pretending there is leads to vanity metrics. There are three families, each answering a different question, and a well-run facility reports from all three because optimizing one in isolation degrades the others. The canonical homes differ by family: the facility metrics are defined and measured in Chapter 15.1, the useful-work metrics (MFU, MBU, goodput, ETTR) in Chapter 14.1 with the reliability reframe in Chapter 12.2, and the unit economics in Chapter 1.8. Here we only orient you to what each family measures and where they conflict.
Facility-efficiency metrics ask: how much of the power and water going in is overhead rather than compute? PUE (total facility energy / IT energy over the same year) is the incumbent; Uptime Institute’s 2026 survey reports a respondent mean and a separate capacity-weighted result (survey published 28 July 2026; Uptime’s analysis 6 August 2026). Keep their weighting distinct; Chapter 15.1 defines both survey benchmarks. Best-in-class liquid-cooled designs reach 1.05-1.15. WUE (litres of water per kWh of IT) ranges from an industry ~1.8-1.9 L/kWh down to ~0 when heat rejection is dry and non-evaporative. ERF/ERE credit reused heat, REF credits renewable energy, and CUE measures carbon per unit of IT energy. PUE and WUE trade against each other, though: evaporative cooling buys a better PUE by spending water (worse WUE), so a facility chasing a headline PUE can quietly become a water glutton. → Chapter 15.4.
Useful-work metrics ask the question PUE cannot: of the energy that did reach the chips, how much produced actual progress? MFU (model FLOPs utilization) and MBU (model bandwidth utilization) measure how close a workload runs to the hardware's theoretical compute and memory-bandwidth ceilings. Goodput — the fraction of wall-clock time spent on useful, non-wasted computation — is the metric that actually governs training economics; a declared 90%-versus-96% pair is useful as a sensitivity scenario, but its endpoints are stipulated rather than drawn from a ClusterMAX or CoreWeave measurement, so they establish neither an industry-average population nor a portable best-in-class baseline. Replace it with the named fleet and job's measured definition, window and event accounting for acceptance or economics. ETTR (effective training time ratio) and its cousin, the time lost to failures and recovery, sit underneath goodput. A facility can post a beautiful 1.08 PUE and still waste a quarter of its energy on a fabric that starves the all-reduce or a checkpoint cadence that loses hours per failure — PUE would never show it. This is the GOODPUT thread of the guide, and it is why facility efficiency is necessary but never sufficient. → Chapter 12.2.
Unit-economics metrics ask: what does the work cost? Tokens-per-joule (and its inverse, joules-per-token) is the rising efficiency metric for inference, because it captures the whole stack — chip, fabric, cooling, software — in one number tied to the product. $/GPU-hr (owned cost requires complete cash inputs; rental price depends on the dated service contract) is the supply-side cost; $/M-tokens (guide-derived self-hosted 70B scenario: $1.90-$2.50/M tokens; Ramp customer cohort: about $10/M -> about $2.50/M in March 2025; not a universal market average) is the demand-side price. These are the numbers the business actually lives on, and they integrate everything above: a facility with great PUE and poor goodput, or great goodput and a bad power contract, shows up here as a bad $/M-tokens. → Chapter 1.8.
| Family | Question it answers | Headline metrics | 2026 reference band | Optimizing this alone breaks |
|---|---|---|---|---|
| Facility efficiency | How much input is overhead, not compute? | PUE, WUE, ERF/ERE, REF, CUE | PUE: 2026 respondent mean / 2026 capacity-weighted result; WUE ~0-1.9 L/kWh | WUE (evaporative cooling) and capex (chasing 1.0x) |
| Useful work | Of energy that reached the chips, how much did real work? | MFU, MBU, goodput, ETTR | Goodput 90% vs 96% sensitivity only; MFU workload-dependent | Cost (over-building fabric/redundancy for marginal goodput) |
| Unit economics | What does the work cost or sell for? | tokens/joule, $/GPU-hr, $/M-tokens | Owned $/GPU-hour from Chapter 1.8’s inputs; $1.90-$2.50/M tokens (guide scenario); Ramp cohort: about $10/M -> about $2.50/M in March 2025 (not a universal market average) | Quality/SLO (cutting cost by degrading latency or accuracy) |
Scope & caveats
Respondent-survey statistic, not an industry-weighted fleet average. The public claim does not provide a population-weighting method; use it as a dated survey benchmark and retain the sample and survey methodology when comparing it with a named facility or fleet.
2025 survey figure (n=681). Uptime's 2026 survey (published 2026-07-28; analysed 2026-08-06) reports two separately defined successors: 1.52 as the annual survey average and 1.36 on a capacity-weighted basis — the gap is the larger-facility advantage, not an improvement in the same population. Neither is a measured average of the world fleet. Same 2026 survey: modal installed-base rack density reached 11 kW (from 9 kW in 2025) — the installed base, not the AI-factory design point.
Scope & caveats
The near-zero endpoint requires dry, non-evaporative heat rejection; loop closure alone does not determine WUE-site.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Scope & caveats
Contested as of Aug 2026: SemiAnalysis (2026-07-05) reports Kyber NVL144 slipping to 2028 on PCB-midplane manufacturability, NVL576 delayed or volume-restricted on CPO maturity, the NVL72x2 copper stopgap cancelled, and CPO-NVSwitch pushed to Feynman; NVIDIA's same-week reply was 'our roadmap remains intact' with no SKU-level rebuttal. Keep 576/1152 as announced figures; do not schedule against them.
Scope & caveats
Scenario output, not a market benchmark; recompute with the selected instance price and workload-specific throughput.
Scope & caveats
Secondary-source whole-facility planning estimate quoted by Savills in May 2024; no disclosed estimating population or method. No universal multiplier follows from Uptime Tier criteria.
Single-source planning heuristic: Savills (May 2024) attributes it to Dgtl Infra, which publishes no sample, geography, density or estimating method. The older ~10–25% inverts a McKinsey 2011 statement (10–20% saving moving Tier IV→III). Uptime Tiers are outcome-based, so no universal cost multiplier follows from the standard; re-estimate against the actual design.
Availability algebra: the nines and how they compose
Availability numbers are abused in infrastructure marketing because the algebra behind them is rarely shown. Get the vocabulary straight here and you can read any "number of nines" claim for what it actually promises.
Availability is A = MTBF / (MTBF + MTTR) — mean operating time between failures over that operating time plus mean restoration downtime in a repairable two-state model. The headline form is "the nines": 99.9% (three nines) is ~8.8 hours of downtime per year; 99.99% (four nines) is ~52 minutes; 99.999% (five nines) is ~5.3 minutes. These arithmetic conversions do not define a facility classification or predict realized service availability. Apply the Uptime outcomes separately: Tier III requires concurrent maintainability; Tier IV adds fault tolerance under the standard's defined conditions. The two levers are independent: you raise availability either by failing less often (higher MTBF) or by recovering faster (lower MTTR) — and in a cluster where restore and replay dominate lost progress, fast checkpoint-restart and lemon-node detection can buy more goodput per dollar than a component-reliability upgrade.
The part that matters for design is composition. Components in series (each one a single point of failure for the path) multiply: a path of ten 99.9% components is 0.999^10 = ~99.0%, not 99.9% — series composition always makes the whole worse than any part. Components in parallel (redundant, where one surviving path suffices) multiply their unavailabilities: two parallel 99% paths give 1 - (0.01 x 0.01) = 99.99%. Both formulas are ideal independent models — the series product and the parallel unavailability product hold only when the elements fail independently — and an N+1 pool is not one-of-two but k-of-n, needing N of the N+1 units, which composes less generously. That is the mathematical case for N+1 and 2N redundancy, and its precondition. A shared controller, a shared cooling loop, or a common-mode failure collapses two "parallel" paths back into series, and the four nines you paid for silently become three. → the redundancy vocabulary (N, N+1, 2N, block- and distributed-redundant, catcher topologies) is defined in Chapter 0.5; the quantitative RBD/Markov/Monte-Carlo machinery is built out in Chapter 12.5. Select against the named service boundary and failure model; Chapter 12.5 owns dependent failures and Chapter 9.4 owns checkpoint recovery.
The three-tier network hierarchy
The last piece of shared vocabulary is the network, because the same words — bandwidth, latency, oversubscription — mean different physics at each of three tiers, and conflating them is how a fabric gets mis-budgeted by an order of magnitude.
Scale-up is the fabric inside a supported load/store and collective-communication domain — NVLink and its peers — connecting the GPUs that act as one large accelerator. It is the fastest, shortest-reach, highest-cost-per-bit tier: ~1.8 TB/s per GPU on NVLink5 (a bidirectional aggregate, i.e. ~900 GB/s each way), a 130 TB/s aggregate inside an NVL72 rack, mostly copper because the reach is sub-metre. Tensor- and expert-parallelism live here. Scale-out is the back-end fabric between domains — InfiniBand or Ethernet/RoCE — connecting racks and pods into a cluster, and its per-NIC rate is generation-specific: 400 Gb/s on ConnectX-7-class nodes (~50 GB/s per direction), 800 Gb/s on the ConnectX-8 NICs shipping with current NVL72-class racks (~100 GB/s per direction). Against NVLink5's ~900 GB/s each way that is a factor of roughly 18 on the older NIC and roughly 9 on the current one, which is exactly why the software prefers to keep tight parallelism inside the scale-up domain and push data- and pipeline-parallel traffic onto the scale-out fabric. That is an optimization preference, not a taxonomy: expert-parallel routing crosses nodes whenever the experts span domains, and some implementations run tensor-parallel traffic there too, so budget the fabric against the job's actual parallelization and traffic matrix. This is the tier where oversubscription is a real choice: published fabrics span 1:1, 2:1–3:1, and reported 7:1 examples, but the project ratio follows the measured traffic matrix, topology, placement, failure headroom, and step-time or tail-latency SLO—not the workload label. Scale-across is optional DCI among buildings or campuses for portfolio connectivity, replication, burst or overflow capacity, service failover, and qualified cross-site computation. It does not require one job to span sites; any cross-site workload must fit the measured latency, bandwidth, algorithm, placement, and correlated-failure budget. A claim about "oversubscription" or "latency" is meaningless until you say which tier it describes. → scale-up in Chapter 8.2; scale-out topology and oversubscription in Chapter 8.5; scale-across multi-campus fabric in Chapter 8.8.
Deep dive: why the metric stack must be read together, never alone
The temptation is to pick one headline number and manage to it. Every era of this industry has a favourite: enterprise IT chased PUE, neoclouds chase $/GPU-hr, training labs chase MFU. Each, optimized alone, breaks something one tier away.
Optimize PUE alone and you reach for evaporative cooling — which buys a 1.1 PUE by spending water, wrecking WUE in a water-stressed region, and does nothing for the goodput being lost to a starved fabric. Optimize goodput alone and you can over-build or under-build: neither a 1:1 fabric nor 2N power follows from the metric by itself. Price the measured bottlenecks and each defined outage state, including transfer interruption, post-event loading, recovery time, and common modes, before choosing either. Optimize $/GPU-hr alone and you cut the redundancy, the health-checking, and the spare capacity that goodput depends on — your supply cost looks great while your delivered cost ($/M-tokens, which integrates goodput) quietly rises. The three families are a system: PUE bounds how much energy reaches the chips, useful-work metrics bound how much of that energy does real work, and unit economics integrate both against price. A defensible operating plan reports one number from each family and watches the conflicts between them. A measurement plan that avoids vanity numbers is built in Chapter 15.1; the economics that integrate them in Chapter 1.8. Use matched intervals and boundaries: annual facility kWh / IT kWh is PUE; capacity is kW at a named state; accepted work is neither installed MW nor energy consumed. Capacity growth, delivered-compute growth and electricity growth require separate baselines.
Deep dive: reading a vendor spec sheet without getting fooled
Vendor and developer claims are usually under-specified rather than dishonest, and the reader supplies the missing assumption, often wrongly. A short checklist, drawn from the vocabulary above, catches most of it.
- A capacity in MW: ask which power — critical IT, total facility, or contracted grid draw — and at what PUE. The three differ by a project-specific amount.
- An equipment rating: ask MW or MVA. Transformers and generators commonly carry kVA/MVA ratings; switchgear is specified by voltage, continuous current, interrupting/withstand duty and the applicable standard; at a 0.9 power factor a 100 MW load is ~111 MVA, and a unit sized to the MW figure is undersized.
- A bandwidth in TB/s or Gb/s: ask which network tier (scale-up, scale-out, scale-across) and whether it is per-GPU, per-NIC, or rack-aggregate, and whether the figure is one-way or a bidirectional aggregate. Scale-up per-GPU and scale-out per-NIC bandwidth span ~18x per direction (NVLink5 ~900 GB/s vs a 400 Gb/s NIC's ~50 GB/s each way); comparing NVLink5's 1.8 TB/s bidirectional aggregate against the NIC's one-way line rate wrongly reads as ~36x.
- An availability in nines: ask what it is the availability of — a single node, the facility power path, or the delivered job — and whether the redundant paths are truly independent or share a common-mode failure.
- An efficiency number: ask which family. A great PUE says nothing about goodput; a great $/GPU-hr says nothing about delivered $/M-tokens.
- A density in kW/rack: ask TDP-summed nameplate or measured workload draw with transients — and at what generation, since the figure is on a steep ramp.
Every one of these is a vocabulary collision waiting to become a procurement error. A single number should never stand without the qualifier that pins its meaning. → numbers provenance and vintage discipline in Chapter 0.2. For a capacity commitment, attach the coincident load calculation and failed-state rating; an annual PUE multiplication is an energy forecast.
How the rest of the guide uses this vocabulary
Nothing here is meant to be the last word — each term has a discipline chapter where it is defined rigorously, measured honestly, and pushed to its edge cases. This chapter is the dictionary you keep open while reading the rest. When Part 3 reorders the siting hierarchy around power, it assumes you read "grid draw" and "MVA" correctly. When Part 5 selects a cooling service, it treats measured kW/rack as one screening input alongside heat split and flux, airflow/inlet limits, water conditions, residual-room rejection, climate, service/redundancy, and the refresh tail—not as the verdict. When Part 8 chooses an oversubscription ratio, it assumes you know which network tier is on the table. When Part 12 argues that AI factories optimize goodput rather than availability, it assumes you can do the availability algebra well enough to see why that is a real distinction. Keep availability and accepted-work attainment as separate acceptance conditions. Higher throughput cannot compensate for missing a contracted recovery or latency limit. The three threads of the guide — POWER-BOUND, GOODPUT, and DENSITY-RAMP — are each, at bottom, a discipline about reading one of these metrics honestly across a four-year build against a moving target.
Cite this chapter
Fehn, J. (2026). Vocabulary, Mental Models & the Metric Stack (Chapter 0.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-0-foundations-and-how-to-use-this-guide/0-3-vocabulary-mental-models-and-the-metric-stack (accessed 2026-09-29).
@misc{aidc-0-3,
author = {Fehn, Jacob},
title = {Vocabulary, Mental Models & the Metric Stack (Chapter 0.3)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-0-foundations-and-how-to-use-this-guide/0-3-vocabulary-mental-models-and-the-metric-stack},
note = {Accessed 2026-09-29}
}