The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 12.2

In this chapter · 6 sections
Term help

The AI-Cluster Reliability Rethink: Goodput vs Facility Availability

An AI cluster earns its return on goodput — the share of bought GPU-hours doing useful work — and for checkpointed training the next redundancy dollar buys more goodput than facility nines; for SLA inference it does not, so price both above the facility states the contract requires.

POWER-BOUNDGOODPUTDENSITY-RAMP

What you'll decide here

  1. Which metric the redundancy spend improves — facility availability, productive training time or ML Productivity Goodput — because a green power-and-cooling record can hide paid GPU-hours that produce no retained work; state each resource and time denominator before connecting it to revenue.
  2. Where redundancy actually lives for your workload: in the facility power chain (2N/Tier-IV), in the silicon and storage (capacitance, hot spares, fast-checkpoint tiers), or in the software (elastic training, request retry) — and therefore what the next dollar buys.
  3. How much continuity the thermal/mechanical path can deliver when a CDU or pump fails — distinguish maintained-flow rejection loss, which can use circulating inventory, from stopped-flow loss that can throttle or trip racks before recovery; qualify each on the selected rack and loop.
  4. Whether your cluster is a grid-reliability problem in its own right — a synchronized multi-hundred-MW load swing the utility now models as a fault — and who pays to flatten it.
  5. The point on the goodput-vs-availability curve where you stop buying facility nines and start buying goodput — the crossover that the Chapter 12.5 model quantifies for your failure environment.
Facility availability and cluster goodput measure different things — you can lose 10% of your compute in a Tier IV hall that never blinks.

For sixty years the data-center industry optimized one number: availability — the fraction of time the facility is energized and cooled, measured over a declared window; Tier certification separately assesses topology outcomes. A Tier III site promises concurrent maintainability; a Tier IV site adds fault tolerance — topology guarantees, not the percentage-uptime folklore still quoted from pre-2009 Uptime documents (Chapter 12.1). That metric was correct for the workload it was built around: enterprise applications and web services where the unit of value is a transaction, an outage is a binary up/down event, and a single rack going dark is a contained, recoverable nuisance. Redundancy — N+1, 2N, block- and distributed-redundant power, dual cooling paths — exists to push that one number toward unity.

An AI factory breaks the assumption underneath the metric. A frontier training job is one tightly-coupled supercomputer running synchronously across tens of thousands of accelerators; a single failed GPU forces a restart-all job back to its last surviving checkpoint when it takes out a required rank; replica-group recovery needs its own state and quorum rules. The facility can be at 100.000% availability — every breaker closed, every CDU pumping — and the cluster can still be throwing away a fifth of the money you spent on it, because the GPUs are idle waiting on a straggler, replaying lost steps, or stalled mid-checkpoint. Goodput — the share of bought GPU-hours that becomes useful work — is the number that governs return, and therefore the one the redundancy budget should be optimizing.

Two metrics, and why they diverge

Define the terms precisely, because the whole rethink lives in the gap between them. Facility availability is a property of the physical plant: the fraction of the observation window in which the declared power-and-cooling conditions hold at the specified load interface; the facility classification is a separate statement about assessed topology and operating outcomes. ML Productivity Goodput is a property of the workload. Google's formulation decomposes it as ML Productivity Goodput = Scheduling Goodput × Runtime Goodput × Program Goodput — resource availability × retained forward-progress time × effective FLOP utilization. Keep three labels apart, because they carry different denominators: Provider Goodput (scheduling × runtime) is the part a contract can commit; ML Productivity Goodput multiplies that by MFU; and Meta's ETTR, or effective training time, is a productive-time observation that excludes MFU entirely. With illustrative Scheduling Goodput of 100%, Runtime Goodput of 90% and Program Goodput (MFU) of 40%, Provider Goodput is 90% and ML Productivity Goodput is 36% — one fleet, two denominators and numbers more than a factor of two apart. Low Program Goodput cuts FLOP productivity while retained-progress time stays unchanged; the contract must say which metric and code obligations it covers. In the scheduling and runtime terms, badput includes accelerator init, JIT compilation, data-loading stalls, checkpoint save and restore, wasted progress replayed after a failure, and infrastructure recovery during restarts (Google Cloud, 2024–2025).

The two numbers diverge because most badput is invisible to the facility. When the building loses power, both availability and goodput drop — they agree. But the dominant losses in a real cluster are not facility outages. Meta's published Llama 3 405B snapshot recorded 419 unplanned interruptions over 54 days on 16,384 H100s — roughly one every three hours — of which the paper attributes approximately 78% to hardware and reports 58.7% as GPU issues, while its table counts and percentages do not reconcile (Meta, 2024). Those are cluster-side interruptions, not evidence of a building outage; the paper does not report facility availability for the snapshot. For the same GPUs, window and eligible-time denominator, facility availability is a ceiling on retained-work time, never a floor: you can keep every required facility service present and still throw away progress below it. That ceiling does not compare one site with a surviving regional fleet or an exclusion-adjusted contract. State the boundaries, then price the lost work; polishing the ceiling while the workload bleeds below it misallocates the redundancy budget.

Redundancy moves into the silicon and the software

The deepest consequence of the goodput reframe is that redundancy migrates out of the facility and into the silicon and the software. In the availability model every resilience dollar bought a redundant power or cooling path. In the goodput model mechanisms above the facility boundary compete for recovered GPU-hours per dollar: multi-tier checkpointing that cuts a training restore from tens of minutes to under a minute; a hot-spare pool with fast health-check and drain, so a failed node is swapped in minutes rather than a fabric re-cable; per-GPU capacitance, rack BBUs and facility BESS that ride through the millisecond-to-second transients the cluster's own load swings cause; and, for serving, replica and region capacity that masks a site event the way a second power path never could. The facility chain still owns the states none of those can bridge — a distribution-path fault that drops a whole hall, a coolant-flow loss that trips every rack on the loop — and it must be sized for them. For a checkpointable fleet whose measured losses are dominated by replay and capacity waits, the checkpoint tier and spare pool move ahead of a second utility path after required facility states close. Reprice the loss still left after each purchase, because two upgrades can recover the same GPU-hours.

Where the next redundancy dollar goes: facility availability vs goodput
Redundancy spendLayerWhat it buysTraining relevanceInference relevanceGoodput leverage
N+1 → 2N facility powerFacilityA second independent path: qualifying path faults and maintenance can retain powerModerate when hall events are rare and recovery is short — a hall outage is one interruption among the job’s other failures; a surviving checkpoint bounds replay, not the wait for usable capacityHigh — an always-on serving site with no second region has no other way to mask the eventWeakest per dollar for checkpointable training; decisive for single-site inference
Multi-tier / async checkpointingStorage + softwareLess blocking, replay and restore time for the qualified checkpoint pathHigh when checkpoint and recovery losses dominateDepends on session, model and external-effect stateStrongest single lever for training goodput
Hot-spare GPU pool + fast health-check/drainSilicon + orchestrationFailed node swapped in minutes, not a fabric re-cableHigh — shrinks recovery time per interruptionModerate — keeps replica count above SLOStrong for both, scales with failure rate
Per-GPU capacitance + rack BBU + facility BESSSilicon + facilityRide-through of transients within each device’s qualified envelope; peak reduction depends on the named platform’s load profile, energy capacity and power-control limitsModerate — prevents transient-induced tripsModerate — protects latency SLO during swingsIndirect — avoids badput from nuisance trips
Redundant CDU / pump / dual-loop cooling continuityThermalCoolant flow held through a pump or CDU fault — the new hall-wide failure modeHigh if flow interruption exceeds the selected rack’s qualified limitDepends on surviving replicas, session state and deadlineStrong for both — the cheapest hall-wide outage to prevent
The recurring comparison in an AI factory. 'Leverage' is goodput recovered per incremental dollar for the named workload; the ranking depends on measured loss and recovery; the amounts come from building the Chapter 12.5 model on your own failure environment.

The table is a spend-allocation guide, and it reads differently for the four fleets an operator actually runs. For frontier synchronous training the ranking is checkpoint tier, spare pool, cooling continuity, transient ride-through, then facility path redundancy — the job already restarts every few hours, so a rare hall event is a marginal addition to a loss the first two rows have bounded; the order holds while measured losses are dominated by replay and capacity waits, and because a required-rank loss interrupts the run either way, compare interruption frequency, durable progress and complete recovery before buying. For smaller checkpointable training the order holds and the facility rows fall further. For frontier rack-scale inference on a single site, flip it: there is no checkpoint to fall back on, the SLA is measured in tail latency, and a hall outage is unmasked unless a second region absorbs it — so cooling continuity and the second power path move to the top. For replicated enterprise inference across regions the fleet masks the site, provided a surviving replica also has weights, keys, session state and sufficient capacity, and the facility rows are bought to the contract's floor and no further. Two things hold across all four: loss of coolant flow is now the cheapest hall-wide outage to prevent, and the contract, not the workload label, sets the floor the facility must certify. Chapter 12.3 selects regional recovery capacity.

torchft illustrates recovery for supported replicated training arrangements. Qualify the actual algorithm and surviving group; its support does not extend automatically to every training stack.

Recovery semantics — choose the surviving unit of work
Workload / retained stateMinimum survivor and replay ruleFailure-domain requirementAcceptance outcome
Restart-all synchronous training: model, optimizer, RNG, data positionAll required ranks restart consistently; uncommitted progress replaysRecoverable checkpoint plus compatible full-job capacityResume the correct step within RTO and declared lost-work limit
Replica-group training: algorithm-specific replicated stateDeclared quorum/group survives; rejoin follows optimizer semanticsPlace surviving groups and state outside the lost domainProve progress and numerical correctness under group loss
Replayable inference: immutable weights and request inputSurviving replica; retry within end-to-end deadlineWeights, routing and authorization available at destinationOne qualified result per logical request
Stateful sessions / agent tools: history, KV cache, side-effect recordRecover session or replay prefix; fence writers and deduplicate effectsDurable session and effect records survive lossNo duplicate external action; latency includes re-prefill
Batch with deadline: input manifest and completion ledgerRequeue only incomplete idempotent workRestart capacity arrives before completion deadlineFull manifest completed by deadline
Deferrable batchReplay accepted; no near-term deadlineBackup and eventual compatible capacityCompletion policy satisfied without reserved hot capacity
Recovery algorithms are selected and tested, not inferred from the word training or inference. torchft supports specified replicated training arrangements; it does not establish support for an arbitrary training stack. Implementation homes: Chapters 9.4, 10.7 and 10.11.

The thermal path: where availability disappears

The most under-appreciated consequence of the density ramp is that coolant flow became a dominant single point of cluster-wide failure, and the loop now operates on two different transient clocks. An air-cooled hall carried enormous thermal inertia: chilled-water volume, the air mass of the room, raised-floor plenum. A CRAH failure can leave time to fail over or intervene when coupled thermal mass holds inlet temperature inside its limit; measure that window for the actual failed component. With flow maintained after rejection loss, coolant and metal buffer the heat input. When flow stops, the distant stored water can no longer carry heat away from the chip: the rack's throttle and trip states arrive on the window the OEM's transient data and a controlled commissioning pump-drop test fix for the specific product — an air-hall runbook cannot supply that window.

This relocates the availability problem. A facility can hold a fault-tolerant power topology and still take the entire cluster down through a coolant-distribution-unit fault, a pump trip, or a control-loop oscillation, when the failed loop exceeds its qualified interruption limit; the assessment record must identify whether the in-rack path is included. The design-basis response is to make coolant continuity a first-class redundancy line: redundant pumps and heat exchangers in the CDU, UPS-backed pump power, isolation that lets one CDU fail without starving the loop, and a commissioning test that drops a pump at full load and proves the ride-through — the posture behind the fleet-wide CDU availability Google has reported since 2020 (Chapter 5.11). Concurrent maintainability — the Tier-III property the industry already values — has to be re-earned in the liquid path: you must be able to pull a pump or service a heat exchanger without dropping the rack. Skimp on cooling-loop redundancy and you have built a cluster whose availability is capped by its weakest pump, no matter how many nines the power chain carries. → Chapter 12.1 sets the topology vocabulary; the DLC continuity engineering is in Chapter 5.4.

Deep dive: the two transient clocks in direct-to-chip cooling

Boundary record. Identify the protected racks, CDU, FWS/TCS endpoints, power feeds, valves and controller supplies. For each operating profile, attach the OEM temperature/flow limits and measured transfer trace. A maintained-flow heat-rejection test and a stopped-flow test are separate records; reserve volume demonstrated in the first does not validate the second.

Specify the automatic response and the safe shutdown state before testing. Compare the complete detection-and-transfer sequence with the qualified limit, including measurement uncertainty. A pump-drop test is an approved commissioning procedure with instrumentation and abort criteria, not an instruction to pull a pump from a live production loop. Follow Chapter 13.5 and Chapter 13.6 for the procedure.

The facility as a grid-reliability problem

The reliability rethink runs in both directions. The cluster's own reliability depends on the facility — but the facility has become a reliability problem for the grid, and that coupling now feeds back into the cluster's design-basis. AI training loads are phase-coherent and synchronized: tens of thousands of GPUs step from idle to peak and back together, every training step, producing load swings of hundreds of megawatts on sub-second timescales. A separate mechanism can drop the entire load: protection response to a grid fault. In a 2024 Virginia event, ~1,500 MW of data-center load tripped off during a six-fault, 82-second sequence on a 230 kV line — enough that the surviving generation had to absorb the imbalance, and enough that NERC issued a rare Level 3 Essential Actions Alert and now treats large data centers as grid actors expected to ride through faults — an expectation the alert recommends but does not yet enforce, with a penalty-backed reliability standard still in development (NERC / Utility Dive, 2026).

Ride-through has become a goodput concern as much as a grid-interconnection one. A cluster that actually loses IT power, cooling or required state during a survivable grid disturbance converts that event into a full restart, paid in badput. A grid-visible transfer to backup while the IT keeps running incurs no such restart loss. The mitigation is the same transient-absorption stack that protects against the cluster's own load swings — per-GPU capacitance, rack BBUs, facility BESS, and NVIDIA's 400 J/GPU Intelligent Power Smoothing, which NVIDIA states can reduce peak current demand by up to 25% — now also tuned to keep the cluster online through utility-side faults rather than dropping load (NVIDIA / SemiAnalysis, 2025–2026). The choice is to engineer the facility to ride through grid disturbances — storage and smoothing capex, plus a regulator-facing study — and test delivered IT power, coolant response and recovery on the same event clock; otherwise a grid disturbance can become a cluster restart paid in goodput, while the utility still sees an unexplained load departure. Qualify the storage over that full envelope. → the full grid-interactive engineering — reactive support, frequency response, ride-through curves at the point of interconnection — is canonical in Chapter 4.10; the storage that backs it in Chapter 4.5.

419 / 54 days
unplanned interruptions on 16,384 H100s (~1 every 3 hr); paper reports ~78% hardware and 58.7% GPU issues; facility availability was not reported
Scope & caveats

Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.

The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

~7 days / one 512-H100 cluster
reported MTBF for one 512-H100 cluster at a top-tier operator
Scope & caveats

SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.

~25–40% facilityestimate
2024 Savills whole-facility heuristic; price required states from the actual scope
Scope & caveats

Secondary-source whole-facility planning estimate quoted by Savills in May 2024; no disclosed estimating population or method. No universal multiplier follows from Uptime Tier criteria.

Single-source planning heuristic: Savills (May 2024) attributes it to Dgtl Infra, which publishes no sample, geography, density or estimating method. The older ~10–25% inverts a McKinsey 2011 statement (10–20% saving moving Tier IV→III). Uptime Tiers are outcome-based, so no universal cost multiplier follows from the standard; re-estimate against the actual design.

+6.59% goodput; save <5 min; restore <1 min
Google’s 35K-chip TPU v5p case: reported ML Goodput uplift and checkpoint-path latencies; restore latency is not complete job recovery time
Scope & caveats

Google reports the figure as a 6.59% increase in ML Goodput on one 35K-chip TPU v5p workload; that wording does not establish 6.59 percentage points absolute. The save and restore latencies are checkpoint-path measurements, not an end-to-end detection-to-resumed-training MTTR for an arbitrary GPU fleet.

~1,500 MW
protection-driven data-center load loss during a six-fault, 82 s grid disturbance on a 230 kV line (VA, Jul 2024); later cited in NERC's Level 3 alert (2026)
Scope & caveats

Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.

The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.

400 J/GPU
NVIDIA Vera Rubin rack-level energy storage normalized per GPU; vendor-stated power-smoothing envelope, not facility ride-through
Scope & caveats

NVIDIA's stated design figure for the Vera Rubin power-shelf PSU capacitor system (NVIDIA Vera Rubin POD, 16 March 2026); the platform entered full production in August 2026. Vendor design statement, not an independent field measurement.

~43.4%
large-LLM job failure rate, top-5% most resource-intensive tasks (Alibaba Unicron); ~37% hardware-attributed, ~73% restart-recoverable
Scope & caveats

Failure rate of the top-5% most resource-intensive large-LLM jobs in one production fleet (Alibaba, Unicron paper), not a per-job, per-GPU or fleet-wide rate. ~37% of the failures are hardware-attributed and ~73% are recoverable by restart on that fleet; the paper does not establish device-population exposure, so this must not be annualized into a component failure rate or applied to another operator's job mix.

Mapping the rethink onto the standards

None of this means the standards are wrong — it means they answer a question that is no longer the binding one. Uptime Tiers, TIA-942 Rated levels, and EN 50600 Availability Classes assess their defined infrastructure requirements; none certifies a numerical goodput guarantee. A Tier IV building tells a tenant the assessed power and cooling paths preserve the required service through the Tier IV fault conditions, with continuous cooling and separately stated outcomes for additional faults during maintenance; it says nothing about whether the cluster inside it loses 10% or 20% of its bought GPU-hours to badput the facility never sees. The standards remain the right tool for the question they answer — they govern the ceiling — but they cannot be the design-basis metric for a workload whose return lives in the gap below that ceiling.

The practical reconciliation is a two-tier design-basis: certify the facility to the availability class the workload's floor requires (higher continuity can pay for an unmasked interactive endpoint, while checkpointable training can price replay — both still owe the required states and customer obligations), then run a separate goodput design-basis that governs the silicon/storage/software redundancy the standards never touch. This is also where the redundancy primer in Chapter 0.5 gets extended: N, N+1, 2N and the distributed-redundant topologies are still the vocabulary, but you now apply them in two places — the facility power/cooling chain and the compute resilience stack — and the goodput model decides which application earns the spend. The standards landscape and topology selection are detailed in Chapter 12.1; the SLA that contracts goodput rather than availability is the subject of Chapter 12.4.

Availability-shaped vs goodput-shaped design basis
AxisFacility-state design basisWorkload-service design basis
Primary metricMeasured facility-state availability; separate classificationEffective accelerator-time (goodput %)
Outage modelBinary up/down at the buildingContinuous badput leakage below a perfect ceiling
Where resilience livesFacility topology — redundant power and cooling pathsCompute stack — checkpoint cadence, hot spares, fast recovery
Power topology inputTier/Class redundancy choice against named maintenance and fault statesThe interruption rate and recovery SLO the workload actually tolerates
Dominant failure to engineer againstUtility outage, switchgear faultGPU/HBM faults, stragglers, slow recovery, cooling loss
Next-dollar priorityClose required physical statesPrice useful-work recovery under the same events
Evidence supplied byUptime / TIA-942 / EN 50600Acceptance and progress records; no goodput certification
Both records apply to training and serving; choose spend from their shared event inventory and service boundary.

The goodput-availability tradeoff curve

Put the mechanisms on one economic curve: incremental lifecycle cost on the x-axis, goodput or realized service value on the y-axis. For a training fleet whose largest avoidable losses are addressed first, the curve starts steep and flattens as those losses disappear; event frequency, recovery time and remaining shared causes set the actual slope and step order. The first dollars above the required facility-state floor go to the largest avoidable loss per dollar. Multi-tier checkpointing attacks wasted progress and slow recovery: Google’s June 16, 2025 report gives a 6.59% ML Goodput increase on a 35K-chip TPU v5p workload and checkpoint restore under one minute, not an end-to-end fleet MTTR guarantee. Spare pools and health checks rise when capacity waits dominate; cooling continuity, transient ride-through and the second facility path compete on the events they prevent. Price each rung against the loss still left after the previous purchase, or pay twice to recover the same GPU-hours. Serving fleets bend the curve differently: request retry and replica, zone or region capacity can mask a site event outright, so a single-site outage is recoverable at the fleet layer and 2N is not automatic; a single-site serving fleet has none of that, and for it the facility rungs climb back to the top. The right budget is where marginal goodput per dollar equalizes across the facility, the compute stack and the fleet — the point Chapter 12.5 locates for a named failure environment.

Deep dive: why facility availability alone does not determine ~90% goodput

Take the Meta Llama 3 405B numbers at face value: 419 interruptions in 54 days on 16,384 GPUs, 78% hardware-caused, and yet over 90% effective training time achieved. Alibaba's Unicron study reports the job-level counterpart on its own production fleet — a 43.4% failure rate among the top-5% most resource-intensive large-LLM jobs, about 37% of it hardware-attributed and roughly 73% recoverable by restart — which is the failure population the recovery machinery is sized against, not a rate to port onto another fleet. Decompose where the other ~10% went, because it shows why facility availability alone cannot explain the number. The losses are: wasted progress — work done since the last checkpoint, thrown away on each interruption (mitigated by checkpoint cadence, the Young/Daly optimal interval); infrastructure recovery — the time to detect the failure, drain the bad node, reschedule, and reload state (mitigated by fast health-checks and multi-tier checkpoint restore); stragglers — the whole synchronous job moving at the speed of its slowest rank, so one degraded 'lemon' GPU taxes thousands of healthy ones (mitigated by lemon-node detection and eviction). Steady-state MFU is not one of the missing ten points: effective training time is measured over productive wall-clock, so MFU multiplies it rather than eating into it. Those buckets describe training interruption and recovery mechanisms; they are not a measurement of facility availability. The paper does not publish rack-power or cooling availability for the snapshot, so it cannot establish what the facility did during every interruption.

The consequence for capital allocation is direct. Paid GPU-hours lost to failures and recovery can be won back through spare capacity, checkpoint-storage bandwidth, health-check tooling and elastic-training engineering, or through facility continuity when those events dominate. Measure the badput buckets before buying: an interruption count alone cannot tell you which investment earns the recovery. The 90%-versus-96% sensitivity is a six-point swing on the same bought GPU-hours — more than any facility nine the building could add above a ceiling it already holds. → the checkpoint math in Chapter 9.4; serving-side goodput in Chapter 10.11.

This chapter is the conceptual hinge of Part 12. The standards and topologies it reframes are detailed in Chapter 12.1, and the redundancy vocabulary it extends into two layers comes from Chapter 0.5. The thermal-continuity engineering that the liquid path now demands is in Chapter 5.4; the storage and grid-interactive behaviour behind ride-through are in Chapter 4.5 and Chapter 4.10. The goodput stack itself — checkpointing in Chapter 9.4 and serving-side goodput-optimal scheduling in Chapter 10.11 — is where the redundancy this chapter redirects actually lives. The crossover point on the tradeoff curve is quantified by the reliability model in Chapter 12.5, fed by the failure rates of Chapter 14.3; the goodput SLA that contracts the result is Chapter 12.4, and the geographic-failover layer above it is Chapter 12.3.

Choose recovery mechanisms that preserve the workload’s required state and complete inside its deadline, then price them against the same facility events. Reserve and storage expenditure earns its place through recovered work; underfunding the required physical states leaves a loss the software cannot mask. Carry the event-to-job mapping from Chapter 8.9 into the model, and carry the resulting measurements to the useful-output ledger in Chapter 14.1.

Cite this chapter
Fehn, J. (2026). The AI-Cluster Reliability Rethink: Goodput vs Facility Availability (Chapter 12.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-2-the-ai-cluster-reliability-rethink-goodput-vs-facility-availability (accessed 2026-09-29).
@misc{aidc-12-2,
  author       = {Fehn, Jacob},
  title        = {The AI-Cluster Reliability Rethink: Goodput vs Facility Availability (Chapter 12.2)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-2-the-ai-cluster-reliability-rethink-goodput-vs-facility-availability},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit