Chapter 12.2
In this chapter · 6 sections
The AI-Cluster Reliability Rethink: Goodput vs Facility Availability
An AI cluster earns its return on goodput — the share of bought GPU-hours doing useful work — and for checkpointed training the next redundancy dollar buys more goodput than facility nines; for SLA inference it does not, so price both above the facility states the contract requires.
What you'll decide here
- Which metric the redundancy spend improves — facility availability, productive training time or ML Productivity Goodput — because a green power-and-cooling record can hide paid GPU-hours that produce no retained work; state each resource and time denominator before connecting it to revenue.
- Where redundancy actually lives for your workload: in the facility power chain (2N/Tier-IV), in the silicon and storage (capacitance, hot spares, fast-checkpoint tiers), or in the software (elastic training, request retry) — and therefore what the next dollar buys.
- How much continuity the thermal/mechanical path can deliver when a CDU or pump fails — distinguish maintained-flow rejection loss, which can use circulating inventory, from stopped-flow loss that can throttle or trip racks before recovery; qualify each on the selected rack and loop.
- Whether your cluster is a grid-reliability problem in its own right — a synchronized multi-hundred-MW load swing the utility now models as a fault — and who pays to flatten it.
- The point on the goodput-vs-availability curve where you stop buying facility nines and start buying goodput — the crossover that the Chapter 12.5 model quantifies for your failure environment.
For sixty years the data-center industry optimized one number: availability — the fraction of time the facility is energized and cooled, measured over a declared window; Tier certification separately assesses topology outcomes. A Tier III site promises concurrent maintainability; a Tier IV site adds fault tolerance — topology guarantees, not the percentage-uptime folklore still quoted from pre-2009 Uptime documents (Chapter 12.1). That metric was correct for the workload it was built around: enterprise applications and web services where the unit of value is a transaction, an outage is a binary up/down event, and a single rack going dark is a contained, recoverable nuisance. Redundancy — N+1, 2N, block- and distributed-redundant power, dual cooling paths — exists to push that one number toward unity.
An AI factory breaks the assumption underneath the metric. A frontier training job is one tightly-coupled supercomputer running synchronously across tens of thousands of accelerators; a single failed GPU forces a restart-all job back to its last surviving checkpoint when it takes out a required rank; replica-group recovery needs its own state and quorum rules. The facility can be at 100.000% availability — every breaker closed, every CDU pumping — and the cluster can still be throwing away a fifth of the money you spent on it, because the GPUs are idle waiting on a straggler, replaying lost steps, or stalled mid-checkpoint. Goodput — the share of bought GPU-hours that becomes useful work — is the number that governs return, and therefore the one the redundancy budget should be optimizing.
Two metrics, and why they diverge
Define the terms precisely, because the whole rethink lives in the gap between them. Facility availability is a property of the physical plant: the fraction of the observation window in which the declared power-and-cooling conditions hold at the specified load interface; the facility classification is a separate statement about assessed topology and operating outcomes. ML Productivity Goodput is a property of the workload. Google's formulation decomposes it as ML Productivity Goodput = Scheduling Goodput × Runtime Goodput × Program Goodput — resource availability × retained forward-progress time × effective FLOP utilization. Keep three labels apart, because they carry different denominators: Provider Goodput (scheduling × runtime) is the part a contract can commit; ML Productivity Goodput multiplies that by MFU; and Meta's ETTR, or effective training time, is a productive-time observation that excludes MFU entirely. With illustrative Scheduling Goodput of 100%, Runtime Goodput of 90% and Program Goodput (MFU) of 40%, Provider Goodput is 90% and ML Productivity Goodput is 36% — one fleet, two denominators and numbers more than a factor of two apart. Low Program Goodput cuts FLOP productivity while retained-progress time stays unchanged; the contract must say which metric and code obligations it covers. In the scheduling and runtime terms, badput includes accelerator init, JIT compilation, data-loading stalls, checkpoint save and restore, wasted progress replayed after a failure, and infrastructure recovery during restarts (Google Cloud, 2024–2025).
The two numbers diverge because most badput is invisible to the facility. When the building loses power, both availability and goodput drop — they agree. But the dominant losses in a real cluster are not facility outages. Meta's published Llama 3 405B snapshot recorded 419 unplanned interruptions over 54 days on 16,384 H100s — roughly one every three hours — of which the paper attributes approximately 78% to hardware and reports 58.7% as GPU issues, while its table counts and percentages do not reconcile (Meta, 2024). Those are cluster-side interruptions, not evidence of a building outage; the paper does not report facility availability for the snapshot. For the same GPUs, window and eligible-time denominator, facility availability is a ceiling on retained-work time, never a floor: you can keep every required facility service present and still throw away progress below it. That ceiling does not compare one site with a surviving regional fleet or an exclusion-adjusted contract. State the boundaries, then price the lost work; polishing the ceiling while the workload bleeds below it misallocates the redundancy budget.
Redundancy moves into the silicon and the software
The deepest consequence of the goodput reframe is that redundancy migrates out of the facility and into the silicon and the software. In the availability model every resilience dollar bought a redundant power or cooling path. In the goodput model mechanisms above the facility boundary compete for recovered GPU-hours per dollar: multi-tier checkpointing that cuts a training restore from tens of minutes to under a minute; a hot-spare pool with fast health-check and drain, so a failed node is swapped in minutes rather than a fabric re-cable; per-GPU capacitance, rack BBUs and facility BESS that ride through the millisecond-to-second transients the cluster's own load swings cause; and, for serving, replica and region capacity that masks a site event the way a second power path never could. The facility chain still owns the states none of those can bridge — a distribution-path fault that drops a whole hall, a coolant-flow loss that trips every rack on the loop — and it must be sized for them. For a checkpointable fleet whose measured losses are dominated by replay and capacity waits, the checkpoint tier and spare pool move ahead of a second utility path after required facility states close. Reprice the loss still left after each purchase, because two upgrades can recover the same GPU-hours.
| Redundancy spend | Layer | What it buys | Training relevance | Inference relevance | Goodput leverage |
|---|---|---|---|---|---|
| N+1 → 2N facility power | Facility | A second independent path: qualifying path faults and maintenance can retain power | Moderate when hall events are rare and recovery is short — a hall outage is one interruption among the job’s other failures; a surviving checkpoint bounds replay, not the wait for usable capacity | High — an always-on serving site with no second region has no other way to mask the event | Weakest per dollar for checkpointable training; decisive for single-site inference |
| Multi-tier / async checkpointing | Storage + software | Less blocking, replay and restore time for the qualified checkpoint path | High when checkpoint and recovery losses dominate | Depends on session, model and external-effect state | Strongest single lever for training goodput |
| Hot-spare GPU pool + fast health-check/drain | Silicon + orchestration | Failed node swapped in minutes, not a fabric re-cable | High — shrinks recovery time per interruption | Moderate — keeps replica count above SLO | Strong for both, scales with failure rate |
| Per-GPU capacitance + rack BBU + facility BESS | Silicon + facility | Ride-through of transients within each device’s qualified envelope; peak reduction depends on the named platform’s load profile, energy capacity and power-control limits | Moderate — prevents transient-induced trips | Moderate — protects latency SLO during swings | Indirect — avoids badput from nuisance trips |
| Redundant CDU / pump / dual-loop cooling continuity | Thermal | Coolant flow held through a pump or CDU fault — the new hall-wide failure mode | High if flow interruption exceeds the selected rack’s qualified limit | Depends on surviving replicas, session state and deadline | Strong for both — the cheapest hall-wide outage to prevent |
The table is a spend-allocation guide, and it reads differently for the four fleets an operator actually runs. For frontier synchronous training the ranking is checkpoint tier, spare pool, cooling continuity, transient ride-through, then facility path redundancy — the job already restarts every few hours, so a rare hall event is a marginal addition to a loss the first two rows have bounded; the order holds while measured losses are dominated by replay and capacity waits, and because a required-rank loss interrupts the run either way, compare interruption frequency, durable progress and complete recovery before buying. For smaller checkpointable training the order holds and the facility rows fall further. For frontier rack-scale inference on a single site, flip it: there is no checkpoint to fall back on, the SLA is measured in tail latency, and a hall outage is unmasked unless a second region absorbs it — so cooling continuity and the second power path move to the top. For replicated enterprise inference across regions the fleet masks the site, provided a surviving replica also has weights, keys, session state and sufficient capacity, and the facility rows are bought to the contract's floor and no further. Two things hold across all four: loss of coolant flow is now the cheapest hall-wide outage to prevent, and the contract, not the workload label, sets the floor the facility must certify. Chapter 12.3 selects regional recovery capacity.
torchft illustrates recovery for supported replicated training arrangements. Qualify the actual algorithm and surviving group; its support does not extend automatically to every training stack.
| Workload / retained state | Minimum survivor and replay rule | Failure-domain requirement | Acceptance outcome |
|---|---|---|---|
| Restart-all synchronous training: model, optimizer, RNG, data position | All required ranks restart consistently; uncommitted progress replays | Recoverable checkpoint plus compatible full-job capacity | Resume the correct step within RTO and declared lost-work limit |
| Replica-group training: algorithm-specific replicated state | Declared quorum/group survives; rejoin follows optimizer semantics | Place surviving groups and state outside the lost domain | Prove progress and numerical correctness under group loss |
| Replayable inference: immutable weights and request input | Surviving replica; retry within end-to-end deadline | Weights, routing and authorization available at destination | One qualified result per logical request |
| Stateful sessions / agent tools: history, KV cache, side-effect record | Recover session or replay prefix; fence writers and deduplicate effects | Durable session and effect records survive loss | No duplicate external action; latency includes re-prefill |
| Batch with deadline: input manifest and completion ledger | Requeue only incomplete idempotent work | Restart capacity arrives before completion deadline | Full manifest completed by deadline |
| Deferrable batch | Replay accepted; no near-term deadline | Backup and eventual compatible capacity | Completion policy satisfied without reserved hot capacity |
The thermal path: where availability disappears
The most under-appreciated consequence of the density ramp is that coolant flow became a dominant single point of cluster-wide failure, and the loop now operates on two different transient clocks. An air-cooled hall carried enormous thermal inertia: chilled-water volume, the air mass of the room, raised-floor plenum. A CRAH failure can leave time to fail over or intervene when coupled thermal mass holds inlet temperature inside its limit; measure that window for the actual failed component. With flow maintained after rejection loss, coolant and metal buffer the heat input. When flow stops, the distant stored water can no longer carry heat away from the chip: the rack's throttle and trip states arrive on the window the OEM's transient data and a controlled commissioning pump-drop test fix for the specific product — an air-hall runbook cannot supply that window.
This relocates the availability problem. A facility can hold a fault-tolerant power topology and still take the entire cluster down through a coolant-distribution-unit fault, a pump trip, or a control-loop oscillation, when the failed loop exceeds its qualified interruption limit; the assessment record must identify whether the in-rack path is included. The design-basis response is to make coolant continuity a first-class redundancy line: redundant pumps and heat exchangers in the CDU, UPS-backed pump power, isolation that lets one CDU fail without starving the loop, and a commissioning test that drops a pump at full load and proves the ride-through — the posture behind the fleet-wide CDU availability Google has reported since 2020 (Chapter 5.11). Concurrent maintainability — the Tier-III property the industry already values — has to be re-earned in the liquid path: you must be able to pull a pump or service a heat exchanger without dropping the rack. Skimp on cooling-loop redundancy and you have built a cluster whose availability is capped by its weakest pump, no matter how many nines the power chain carries. → Chapter 12.1 sets the topology vocabulary; the DLC continuity engineering is in Chapter 5.4.
Deep dive: the two transient clocks in direct-to-chip cooling
Boundary record. Identify the protected racks, CDU, FWS/TCS endpoints, power feeds, valves and controller supplies. For each operating profile, attach the OEM temperature/flow limits and measured transfer trace. A maintained-flow heat-rejection test and a stopped-flow test are separate records; reserve volume demonstrated in the first does not validate the second.
Specify the automatic response and the safe shutdown state before testing. Compare the complete detection-and-transfer sequence with the qualified limit, including measurement uncertainty. A pump-drop test is an approved commissioning procedure with instrumentation and abort criteria, not an instruction to pull a pump from a live production loop. Follow Chapter 13.5 and Chapter 13.6 for the procedure.
The facility as a grid-reliability problem
The reliability rethink runs in both directions. The cluster's own reliability depends on the facility — but the facility has become a reliability problem for the grid, and that coupling now feeds back into the cluster's design-basis. AI training loads are phase-coherent and synchronized: tens of thousands of GPUs step from idle to peak and back together, every training step, producing load swings of hundreds of megawatts on sub-second timescales. A separate mechanism can drop the entire load: protection response to a grid fault. In a 2024 Virginia event, ~1,500 MW of data-center load tripped off during a six-fault, 82-second sequence on a 230 kV line — enough that the surviving generation had to absorb the imbalance, and enough that NERC issued a rare Level 3 Essential Actions Alert and now treats large data centers as grid actors expected to ride through faults — an expectation the alert recommends but does not yet enforce, with a penalty-backed reliability standard still in development (NERC / Utility Dive, 2026).
Ride-through has become a goodput concern as much as a grid-interconnection one. A cluster that actually loses IT power, cooling or required state during a survivable grid disturbance converts that event into a full restart, paid in badput. A grid-visible transfer to backup while the IT keeps running incurs no such restart loss. The mitigation is the same transient-absorption stack that protects against the cluster's own load swings — per-GPU capacitance, rack BBUs, facility BESS, and NVIDIA's 400 J/GPU Intelligent Power Smoothing, which NVIDIA states can reduce peak current demand by up to 25% — now also tuned to keep the cluster online through utility-side faults rather than dropping load (NVIDIA / SemiAnalysis, 2025–2026). The choice is to engineer the facility to ride through grid disturbances — storage and smoothing capex, plus a regulator-facing study — and test delivered IT power, coolant response and recovery on the same event clock; otherwise a grid disturbance can become a cluster restart paid in goodput, while the utility still sees an unexplained load departure. Qualify the storage over that full envelope. → the full grid-interactive engineering — reactive support, frequency response, ride-through curves at the point of interconnection — is canonical in Chapter 4.10; the storage that backs it in Chapter 4.5.
Scope & caveats
Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.
The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Scope & caveats
Secondary-source whole-facility planning estimate quoted by Savills in May 2024; no disclosed estimating population or method. No universal multiplier follows from Uptime Tier criteria.
Single-source planning heuristic: Savills (May 2024) attributes it to Dgtl Infra, which publishes no sample, geography, density or estimating method. The older ~10–25% inverts a McKinsey 2011 statement (10–20% saving moving Tier IV→III). Uptime Tiers are outcome-based, so no universal cost multiplier follows from the standard; re-estimate against the actual design.
Scope & caveats
Google reports the figure as a 6.59% increase in ML Goodput on one 35K-chip TPU v5p workload; that wording does not establish 6.59 percentage points absolute. The save and restore latencies are checkpoint-path measurements, not an end-to-end detection-to-resumed-training MTTR for an arbitrary GPU fleet.
Scope & caveats
Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.
The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.
Scope & caveats
NVIDIA's stated design figure for the Vera Rubin power-shelf PSU capacitor system (NVIDIA Vera Rubin POD, 16 March 2026); the platform entered full production in August 2026. Vendor design statement, not an independent field measurement.
Scope & caveats
Failure rate of the top-5% most resource-intensive large-LLM jobs in one production fleet (Alibaba, Unicron paper), not a per-job, per-GPU or fleet-wide rate. ~37% of the failures are hardware-attributed and ~73% are recoverable by restart on that fleet; the paper does not establish device-population exposure, so this must not be annualized into a component failure rate or applied to another operator's job mix.
Mapping the rethink onto the standards
None of this means the standards are wrong — it means they answer a question that is no longer the binding one. Uptime Tiers, TIA-942 Rated levels, and EN 50600 Availability Classes assess their defined infrastructure requirements; none certifies a numerical goodput guarantee. A Tier IV building tells a tenant the assessed power and cooling paths preserve the required service through the Tier IV fault conditions, with continuous cooling and separately stated outcomes for additional faults during maintenance; it says nothing about whether the cluster inside it loses 10% or 20% of its bought GPU-hours to badput the facility never sees. The standards remain the right tool for the question they answer — they govern the ceiling — but they cannot be the design-basis metric for a workload whose return lives in the gap below that ceiling.
The practical reconciliation is a two-tier design-basis: certify the facility to the availability class the workload's floor requires (higher continuity can pay for an unmasked interactive endpoint, while checkpointable training can price replay — both still owe the required states and customer obligations), then run a separate goodput design-basis that governs the silicon/storage/software redundancy the standards never touch. This is also where the redundancy primer in Chapter 0.5 gets extended: N, N+1, 2N and the distributed-redundant topologies are still the vocabulary, but you now apply them in two places — the facility power/cooling chain and the compute resilience stack — and the goodput model decides which application earns the spend. The standards landscape and topology selection are detailed in Chapter 12.1; the SLA that contracts goodput rather than availability is the subject of Chapter 12.4.
| Axis | Facility-state design basis | Workload-service design basis |
|---|---|---|
| Primary metric | Measured facility-state availability; separate classification | Effective accelerator-time (goodput %) |
| Outage model | Binary up/down at the building | Continuous badput leakage below a perfect ceiling |
| Where resilience lives | Facility topology — redundant power and cooling paths | Compute stack — checkpoint cadence, hot spares, fast recovery |
| Power topology input | Tier/Class redundancy choice against named maintenance and fault states | The interruption rate and recovery SLO the workload actually tolerates |
| Dominant failure to engineer against | Utility outage, switchgear fault | GPU/HBM faults, stragglers, slow recovery, cooling loss |
| Next-dollar priority | Close required physical states | Price useful-work recovery under the same events |
| Evidence supplied by | Uptime / TIA-942 / EN 50600 | Acceptance and progress records; no goodput certification |
The goodput-availability tradeoff curve
Put the mechanisms on one economic curve: incremental lifecycle cost on the x-axis, goodput or realized service value on the y-axis. For a training fleet whose largest avoidable losses are addressed first, the curve starts steep and flattens as those losses disappear; event frequency, recovery time and remaining shared causes set the actual slope and step order. The first dollars above the required facility-state floor go to the largest avoidable loss per dollar. Multi-tier checkpointing attacks wasted progress and slow recovery: Google’s June 16, 2025 report gives a 6.59% ML Goodput increase on a 35K-chip TPU v5p workload and checkpoint restore under one minute, not an end-to-end fleet MTTR guarantee. Spare pools and health checks rise when capacity waits dominate; cooling continuity, transient ride-through and the second facility path compete on the events they prevent. Price each rung against the loss still left after the previous purchase, or pay twice to recover the same GPU-hours. Serving fleets bend the curve differently: request retry and replica, zone or region capacity can mask a site event outright, so a single-site outage is recoverable at the fleet layer and 2N is not automatic; a single-site serving fleet has none of that, and for it the facility rungs climb back to the top. The right budget is where marginal goodput per dollar equalizes across the facility, the compute stack and the fleet — the point Chapter 12.5 locates for a named failure environment.
Deep dive: why facility availability alone does not determine ~90% goodput
Take the Meta Llama 3 405B numbers at face value: 419 interruptions in 54 days on 16,384 GPUs, 78% hardware-caused, and yet over 90% effective training time achieved. Alibaba's Unicron study reports the job-level counterpart on its own production fleet — a 43.4% failure rate among the top-5% most resource-intensive large-LLM jobs, about 37% of it hardware-attributed and roughly 73% recoverable by restart — which is the failure population the recovery machinery is sized against, not a rate to port onto another fleet. Decompose where the other ~10% went, because it shows why facility availability alone cannot explain the number. The losses are: wasted progress — work done since the last checkpoint, thrown away on each interruption (mitigated by checkpoint cadence, the Young/Daly optimal interval); infrastructure recovery — the time to detect the failure, drain the bad node, reschedule, and reload state (mitigated by fast health-checks and multi-tier checkpoint restore); stragglers — the whole synchronous job moving at the speed of its slowest rank, so one degraded 'lemon' GPU taxes thousands of healthy ones (mitigated by lemon-node detection and eviction). Steady-state MFU is not one of the missing ten points: effective training time is measured over productive wall-clock, so MFU multiplies it rather than eating into it. Those buckets describe training interruption and recovery mechanisms; they are not a measurement of facility availability. The paper does not publish rack-power or cooling availability for the snapshot, so it cannot establish what the facility did during every interruption.
The consequence for capital allocation is direct. Paid GPU-hours lost to failures and recovery can be won back through spare capacity, checkpoint-storage bandwidth, health-check tooling and elastic-training engineering, or through facility continuity when those events dominate. Measure the badput buckets before buying: an interruption count alone cannot tell you which investment earns the recovery. The 90%-versus-96% sensitivity is a six-point swing on the same bought GPU-hours — more than any facility nine the building could add above a ceiling it already holds. → the checkpoint math in Chapter 9.4; serving-side goodput in Chapter 10.11.
Choose recovery mechanisms that preserve the workload’s required state and complete inside its deadline, then price them against the same facility events. Reserve and storage expenditure earns its place through recovered work; underfunding the required physical states leaves a loss the software cannot mask. Carry the event-to-job mapping from Chapter 8.9 into the model, and carry the resulting measurements to the useful-output ledger in Chapter 14.1.
Cite this chapter
Fehn, J. (2026). The AI-Cluster Reliability Rethink: Goodput vs Facility Availability (Chapter 12.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-2-the-ai-cluster-reliability-rethink-goodput-vs-facility-availability (accessed 2026-09-29).
@misc{aidc-12-2,
author = {Fehn, Jacob},
title = {The AI-Cluster Reliability Rethink: Goodput vs Facility Availability (Chapter 12.2)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-2-the-ai-cluster-reliability-rethink-goodput-vs-facility-availability},
note = {Accessed 2026-09-29}
}