The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 10.7

In this chapter · 7 sections
Term help

Fleet Reliability, Fault Tolerance & Autonomous Recovery

Job-level MTBF collapses with node count, so training reliability is a control-plane problem: detect, fence the failed rank, select a complete surviving checkpoint and verify the restarted step; measure lost progress and the full recovery interval, because every minute of it is goodput.

GOODPUTDENSITY-RAMPPOWER-BOUND

What you'll decide here

  1. Whether your reliability target is facility availability (the legacy 'nines') or goodput/ETTR — for a checkpointable training fleet the two diverge sharply, and optimizing the wrong one buys redundancy the workload does not value (canonical rethink in Chapter 12.2).
  2. Price hot spares and multi-tier checkpointing against the measured interruption and recovery budget; neither is automatically the best goodput purchase.
  3. Compare hot spares, elastic restart and redundant execution including the framework coupling each introduces.
  4. How autonomous the recovery loop is allowed to be — auto-drain and auto-restart on a confidence threshold versus human-gated remediation — and the blast-radius guardrails (rate limits, quarantine, lemon-node ejection) that keep an autonomous loop from amplifying a fault.
  5. The checkpoint cadence and restart-overhead target derived from the named job's measured interruption distribution, save cost, restart cost, and goodput objective (Young/Daly math canonical in Chapter 9.4) — not from accelerator count alone.

A larger synchronous job participates in more component and service failure domains, so job interruption becomes a first-order goodput constraint. Meta reported about 47.7 days MTTF for an 8-GPU job and 7.9 hours for a 1,024-GPU job in its named research-cluster population. Those observations support the direction of scale, but they do not justify a universal six-figure-GPU cadence without fleet telemetry and explicit independence/common-mode, job-membership, software, and event-definition assumptions. That is why fleet reliability is an engineering discipline of its own rather than a facility-uptime line item.

This chapter is about the system that turns that failure rate from a goodput catastrophe into a manageable tax. It is built from three moving parts that must be designed together: the detection-to-recovery loop (how fast you notice, isolate, and restart), the fault-tolerance strategy (hot spares versus elastic/redundant training versus algorithmic resilience), and the checkpoint/restore substrate that decides how much work a failure erases. Every reliability dollar can go to facility nines, to recovery speed, or to spare capacity, and the three buy very different amounts of goodput. We name the canonical homes for the supporting math — checkpoint intervals in Chapter 9.4, the failure taxonomy and fleet AFR data in Chapter 14.3, and the availability-vs-goodput rethink in Chapter 12.2 — and concentrate here on the operational loop that ties them together.

The reliability problem at scale

Two facts collide to produce the modern training-reliability problem. First, synchronous coupling makes every node a single point of failure: in a data-/tensor-/pipeline-parallel run the job advances at the speed of its slowest rank, and a dead rank halts all of them. Second, the number of participating failure domains grows with job size. Under an explicit independent-identical-rate model the aggregate hazard grows with membership, but real fleets also carry correlated hardware, shared-service, network, software, and detection effects. The operational input is therefore the measured effective job-level interruption distribution, not a universal hourly cadence.

Several public empirical anchors are useful but not directly interchangeable. Meta's Llama 3 405B run logged 419 unplanned interruptions over 54 days on 16,384 H100s — about one every three hours — of which approximately 78% were hardware-caused; the paper reports 58.7% as GPU issues although Table 5's counts and percentages do not reconcile, yet the team still achieved over 90% effective training time in that Llama 3 campaign through aggressive automation and only three manual interventions. SemiAnalysis's October 2024 operator playbook reports about 7 days MTBF for one 512-H100 cluster at a top-tier operator and separately recommends at least 3–4 weeks of factory high-temperature burn-in before deployment. The first is not a per-GPU rate; the second is not an observed universal settling time. Alibaba's Unicron production study found a 43.4% failure rate among the top 5% most resource-intensive large jobs, about 37% hardware-attributed and roughly 73% recoverable via restart. At scale, failure is the steady state to engineer for.

The consequence for design is that a facility-topology classification does not measure training output. Uptime Tier outcomes describe topology under defined conditions; a 100k-GPU synchronous job loses far more useful time to internal hardware faults that the facility's 2N power and cooling do nothing to prevent. The metric that governs the return on a training cluster is goodput (equivalently ETTR, effective-training-time ratio): productive GPU-time divided by wall-clock GPU-time. This is the canonical pivot of Chapter 12.2; here it is the lens through which every recovery decision is scored.

The detection-to-recovery loop

Every interruption runs through the same five-stage loop, and the time spent in each stage is what you actually control. Detect the fault; drain the affected node or rack out of the job; diagnose the root cause; remediate (reboot, reseat, RMA, or replace); and restart the job from the last good checkpoint onto healthy hardware. The cluster's goodput is set by how fast this loop closes and how often it has to run. The failure taxonomy that the loop must classify — hard faults, transient faults, and silent data corruption — is canonical in Chapter 14.3; here we treat detection as a given input and focus on the loop's economics.

The non-obvious lever is that detection latency and restart latency dominate, not repair latency. Repair (an RMA, a reseat) happens asynchronously on a drained node while the job runs on a spare; it is off the critical path. What sits on the critical path is the time to notice the fault (a hung collective can stall a job for minutes before a watchdog fires) plus the time to load a checkpoint and re-establish the fabric. This is why the reliability spend with the largest return is rarely 'better hardware': it is faster watchdogs, faster checkpoint loading, and a warm spare ready to slot in. Multi-tier checkpointing has driven restart from the legacy 15–30 minutes down toward under two minutes, and that single change can move goodput by several points at frontier scale.

The detection-to-recovery loop: stages, levers, and what they cost
StageTypical latencyOn critical path?Primary leverFailure mode if neglected
DetectSeconds to several minutesYes — job is stalled while undetectedHeartbeats, collective watchdogs, health checks, SDC scannersA hung rank silently burns GPU-hours until a timeout fires
Drain / isolateSeconds to ~1 minYesTopology-aware eject; quarantine the node/rack from the schedulerFaulty node rejoins and re-fails; flapping job
DiagnoseMinutes to hoursNo — runs on drained nodeAutomated triage workflows; XID/SXID classification; burn-in re-testMis-triage RMAs healthy parts or returns a lemon to service
RemediateMinutes (reboot) to days (RMA)No — off-line on a spareReboot/reseat/reflash; RMA logistics; lemon-node ejectionRepeat-offender 'lemon' nodes silently cap fleet goodput
Restart~1 min restore from a node-local tier vs tens of minutes across the storage fabric — plus relaunch and collective re-initYes — all GPUs idle until resumedMulti-tier / in-memory checkpoint; hot spare; fast fabric re-initA slow restore multiplies every failure into a large goodput loss
Stage latencies are 2024–2026 practitioner ranges for frontier synchronous training; 'on critical path' indicates whether the stage stalls the running job. Figures synthesize Meta (Llama 3 / Revisiting Reliability), Google Cloud multi-tier checkpointing, and NVIDIA Mission Control.

The table is a budget allocator. The three rows marked 'on critical path' — detect, drain, restart — are where wall-clock goodput is won or lost; the two marked off-path can be slow and asynchronous after safe fencing and replacement placement, provided you have compatible spare capacity to keep the job running while they complete. This is the structural argument for hot spares: they convert remediate from an on-path stall into an off-path background task. It is also the argument for multi-tier checkpointing: it attacks restart, a potentially expensive on-path stage, by keeping a recent checkpoint in node-local memory or NVMe provided that copy survives the named failure domain and contains validated recoverable state. The checkpoint cadence and tiering math that governs how much a failure erases is canonical in Chapter 9.4.

Illustrative — stated assumptions. The assumed left state has a replacement that passes software, placement, capacity and acceptance requirements; the right has no eligible replacement. Diagnosis and repair occur after fencing and rejoin allocation only after the acceptance test procedure (ATP). Arrows show dependencies, not durations. Buy qualified reserve when allocation is the controlling delay; shortening a different stage leaves that wait. Checkpoint age and lost progress use Chapter 9.4; this drawing does not rederive the checkpoint interval or imply repair restores a job by itself.

Autonomous hardware recovery: closing the loop without a human

At a fleet failing more than once an hour, a human-in-the-loop recovery process is a bottleneck — the operator becomes the MTTR. The 2025–2026 answer is to make the loop autonomous: a fleet control plane that detects a drained or unhealthy node, runs automated triage to classify the fault, attempts remediation (power-cycle, reflash, re-test) without a ticket, and only escalates to a human when it cannot resolve the fault itself. NVIDIA's Mission Control packages this for GB200/GB300 NVL72 as three coupled components — autonomous job recovery, autonomous hardware recovery, and the NVIDIA Resiliency Extension (NVRx) — running automated health checks at the tray, rack, and system level and executing break-fix workflows that open support tickets only for what cannot auto-resolve. Hyperscalers run their own equivalents (Meta's automation took Llama 3 to over 90% effective training time with only three manual interventions across 54 days), and neocloud operators differentiate on the maturity of exactly this loop.

The decision here is not whether to automate detection — everyone does — but how much authority to grant the loop to act. Full autonomy (auto-drain and auto-restart the job on a confidence threshold) maximizes goodput but can amplify a fault: a mis-classifying triage routine can eject healthy nodes, a restart storm can thrash the scheduler, and a shared firmware bug can be 'remediated' onto every node in turn. The guardrails that make autonomy safe are the same ones that make any control loop safe — rate limits on automated actions, quarantine of repeat offenders (lemon-node ejection), and a circuit breaker that hands control to a human when the action rate or failure rate spikes. An under-automated loop loses goodput to human latency; an over-automated loop without guardrails loses goodput to self-inflicted instability.

Fault tolerance: hot spares vs elastic/redundant training vs algorithmic resilience

Once detection and recovery are fast, the next fork is how the job survives the failure — and there are three families, each trading capacity, framework coupling, and recovery speed differently. The choice determines both the spare-capacity tax you pay and how deeply reliability is wired into the training stack.

Hot spares + fast restart is the operationally simplest and most common posture: hold a pool of healthy GPUs idle (typically a few percent of the fleet), and on failure drain the bad node, slot in a spare, and restart the job from the last checkpoint. The cost is the idle spare capacity and the full restart latency; the virtue is that it is framework-agnostic and easy to reason about. Elastic / redundant training lets the job continue on fewer nodes (shrinking the world size and re-sharding) or run with redundant replicas so a single failure does not require a full restart — eliminating the spare tax and much of the restart stall, at the price of coupling reliability tightly to the training framework (the scheduler, the parallelism plan, and the collective library must all cooperate to reshard live). Algorithmic fault tolerance goes further into the math: techniques like nonuniform/elastic tensor parallelism reduce the goodput amplification a single GPU failure causes, and redundant-computation or erasure-coded schemes let the job tolerate a fault without re-execution — the lowest overhead in principle, but the least mature and the most workload-specific. A fourth family emerged commercially in 2026: live GPU migration (Clockwork TorchPass), which sits between the framework and the scheduler and migrates a running job onto a spare without a restart — vendor-benchmarked at zero lost steps on a 64-H200 failure injection versus 869 for checkpoint-restart, and independently tested by SemiAnalysis, which rated it the fastest fault-tolerant approach it measured. It is licensed software, the spare pool is still required, and the hard-failure path degrades toward just-in-time checkpointing — so treat it as a restart-latency eliminator for planned and XID-class faults, not a repeal of the checkpoint discipline. These map onto the operational reliability-engineering treatment in Chapter 14.4.

Fault-tolerance strategies for synchronous training
StrategyHow it survives a failureSpare-capacity taxRecovery latencyFramework couplingBest fit
Hot spares + fast restartSwap in a healthy spare; restart from checkpoint~2–5% of fleet held idleRestart-bound (node-local restore ~1 min; tens of minutes from storage)Low — scheduler-level, framework-agnosticDefault for most operators; large stable runs
Elastic trainingShrink world size / re-shard onto survivors; resumeNone (no idle reserve)Re-shard + resume; no spare provisioning waitHigh — scheduler + parallelism + collectives must reshard liveLong runs where spare capacity is scarce or costly
Redundant trainingRedundant replicas absorb the loss; no full restartReplica overhead (compute, not idle)Near-zero stall on a single failureHigh — requires replicated execution planHighest-value runs where any stall is unacceptable
Algorithmic fault toleranceNonuniform/elastic TP or coded redundancy bounds the lossLow (math, not idle capacity)Often no restart for the tolerated fault classHighest — baked into the parallelism/algorithmFrontier teams co-designing model, parallelism, and resilience
A practitioner comparison of the three families as of 2026. 'Spare tax' is idle capacity reserved purely for failover; 'framework coupling' is how tightly the strategy ties into the training stack.

Moving down this table changes operational complexity, reserve capacity and recovery speed; elastic or redundant training also pays in framework coupling. Hot spares are something an operations team can run with a stock framework; elastic and redundant training require the training stack itself to be reliability-aware, which is why they are most common at the labs that own their framework end-to-end. The right answer is a function of how scarce your spare capacity is (a power-bound fleet must price every idle reserved GPU), how much a stall costs (a contracted run with a deadline cannot tolerate restart storms), and how much control you have over the training stack. A hybrid can use hot spares first and qualified elastic shrink when they are exhausted. A spare must match memory, tuple and fabric locality; free inventory in the wrong domain cannot back the promise. Test simultaneous claims on shared reserve before selling it twice.

7.9 hr
mean-time-to-failure of a 1,024-GPU job vs 47.7 days for an 8-GPU job — named Meta research-cluster observations, not a universal scaling curve
6.50 / 1000
failures per thousand node-days on Meta's RSC-1 cluster (11 months, ~80%+ utilization)
Scope & caveats

Job failures on Meta's A100 RSC-1 research cluster over an 11-month window at ~80%+ utilization, divided by allocated node-days (job runtime x allocated nodes). An event-per-exposure rate for one named cluster and workload mix, not unique failed-FRU counts and not a portable component failure intensity; the companion RSC-2 figure is 2.34 on the same basis.

419 / 54 days
unplanned interruptions on 16,384 H100s during Llama 3 405B (~1 every 3 hr); paper reports ~78% hardware and 58.7% GPU issues, but Table 5 counts and percentages do not reconcile
Scope & caveats

Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.

The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.

~7 days / one 512-H100 cluster
reported MTBF for one 512-H100 cluster at a top-tier operator
Scope & caveats

SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.

~43.4%
large-LLM-job failure rate, top-5% most resource-intensive tasks (Alibaba Unicron); ~37% hardware-attributed; ~73% recoverable via restart
Scope & caveats

Failure rate of the top-5% most resource-intensive large-LLM jobs in one production fleet (Alibaba, Unicron paper), not a per-job, per-GPU or fleet-wide rate. ~37% of the failures are hardware-attributed and ~73% are recoverable by restart on that fleet; the paper does not establish device-population exposure, so this must not be annualized into a component failure rate or applied to another operator's job mix.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

+6.59% goodput; save <5 min; restore <1 min
Google’s named TPU checkpoint save/restore result; end-to-end recovery has additional stages
Scope & caveats

Google reports the figure as a 6.59% increase in ML Goodput on one 35K-chip TPU v5p workload; that wording does not establish 6.59 percentage points absolute. The save and restore latencies are checkpoint-path measurements, not an end-to-end detection-to-resumed-training MTTR for an arbitrary GPU fleet.

Sizing the recovery: where the next reliability dollar goes

The reliability budget has three competing claims — raise MTBF (better hardware, more burn-in), shrink MTTR (faster detect/restart, hot spares), or add facility nines (2N power, redundant cooling) — and at frontier scale they are not equally productive. The goodput of a checkpointable job is, to first order, uptime fraction = MTBF / (MTBF + MTTR + lost-work-per-failure), where lost work is modeled from the interval between checkpoints durably committed inside the declared failure domain; the RPO boundary is the newest checkpoint recoverable in that domain. This is a failure-loss illustration rather than a full goodput accounting: it carries no term for the ordinary cost of checkpointing a healthy job, so it tends to 100% in the no-failure limit even where every 100 s of useful training pays 10 s of blocking save — a true 90.9%. Take the complete accounting — productive work, ordinary saves, detection and recovery, recomputation — from Chapter 9.4. Because MTBF falls as you scale and is hard to move (the hardware is what it is), the dominant levers are MTTR (recovery speed) and lost-work-per-failure (checkpoint cadence) — both of which are software and control-plane investments, not facility ones.

At a named fleet’s measured failure distribution, checkpointing and hot spares may dominate the next-dollar goodput comparison because facility topology does not prevent silicon, software, or membership failures. That does not make 2N categorically wasteful: it prevents only the modeled facility states, whose frequency, interruption, post-event load, recovery consequence, common modes, and contractual value must be compared with compute-stack resilience. The Young/Daly checkpoint derivation is in Chapter 9.4; the state-based economic comparison is in Chapter 12.5.

Deep dive: derive checkpoint cadence from the job’s interruption exposure

Chapter 9.4 derives checkpoint cadence from measured whole-job interruptions and checkpoint/restart cost. For a synchronous job exposed to each member failure, adding members can shorten interruption intervals; correlated rack, fabric and software failures require a separate model. Meta’s cited 1,024-GPU result is not a forecast for a 100k-GPU job. Chapter 9.4 distinguishes unique saved state, capture, drain and durable commit; neither <14 bytes/parameter nor a universal overlap target describes every checkpoint. A recent node-local copy speeds restart only for failure domains it survives.

The concrete target must come from the named job's measured effective interruption distribution. Feed that distribution, checkpoint blocking cost, restart cost, and ETTR objective into the model; then co-design cadence, tiering, and autonomous recovery around the result. A two-minute target from an unstated six-figure-GPU independence extrapolation is not a portable operating requirement.

Inference fleets: a different reliability problem

Everything above is the training story, where one fault stalls one tightly-coupled job. An online-inference fleet inverts almost every assumption and therefore inverts the reliability strategy. At the API, inference requests remain retry-oriented, but the infrastructure has two serving tiers: in independently replicated serving, a node failure drops one replica's in-flight requests; in frontier serving, one replica can span an NVL72 scale-up domain, so a node or fabric failure can remove the distributed replica and every request it is serving. There is no training checkpoint to restore — the unit of failure is a request at the API and a replica or domain in the infrastructure. The reliability targets are correspondingly different: request success rate and SLO attainment (TTFT/TPOT within budget) rather than goodput-as-effective-training-time. Required maintenance and fault outcomes select the facility topology, not the inference label; revenue-critical serving often justifies a stronger posture than checkpointable batch work.

So the fault-tolerance toolkit shifts. For inference you invest in fast health-checking and load-balancer ejection (pull a sick replica out of rotation in seconds) and, for distributed replicas, domain-aware health, replica-level draining, and whole-domain spare/placement policy, geographic and zonal redundancy (so a rack, hall, or region failure degrades capacity rather than availability), and graceful degradation (shed or queue low-priority traffic, fall back to a smaller model, preserve the SLO for what remains) rather than checkpoint cadence and hot training spares. KV-cache recovery cost follows context length and replica design: short prompts are cheap to recompute, while long-context prefill can make recomputation or replication a material capacity and latency decision. Inference recovery therefore couples API retry with replica/domain recovery — and prefix-cache locality (Chapter 9.7) means losing a replica still costs cache warmth and latency. The serving-engine and SLO mechanics that this reliability posture protects are owned in Chapter 10.11; the multi-region failover and DR design in Chapter 12.3; and the SLA/goodput-contract framing in Chapter 12.4.

Reliability posture: training fleet vs inference fleet
DimensionSynchronous training fleetOnline inference fleet
Unit of failureThe whole job (one node stalls all ranks)Request at the API; replica/domain in the infrastructure
Headline metricGoodput / ETTR (effective training time)Request success rate + SLO attainment (TTFT/TPOT)
Recovery primitiveDrain → restart from checkpointEject affected replica/domain from service → retry requests elsewhere
Facility topology inputsNamed maintenance/fault states, transfer interruption, post-event capacity, recovery SLO, independence and common modesNamed maintenance/fault states, transfer interruption, remaining replica/failover capacity, recovery SLO, independence and common modes
Spare strategyHot GPU spares / elastic reshardOver-provisioned replicas/domains + zonal/geo redundancy
Dominant leverCheckpoint cadence + restart latencyHealth-check speed + domain-aware placement + graceful degradation
The two archetypes optimize different reliability targets and therefore spend the reliability budget on different mechanisms. Mapped to the workload archetypes of Chapter 1.1.
Deep dive: silent data corruption — the failure the recovery loop can't see

Every mechanism in this chapter assumes the fault announces itself: a node hangs, a NIC drops, a health check goes red, and the loop fires. Silent data corruption (SDC) is the failure class that breaks that assumption — a GPU that computes the wrong answer and keeps running, returning corrupted gradients that poison the model without tripping any watchdog. At fleet scale SDC is not theoretical: hyperscalers report it as a real and growing contributor, and an undetected SDC event can waste days of training before a loss spike or a downstream eval reveals that the weights are bad — potentially a larger goodput loss than a promptly recovered hard fault, because it corrupts work that looked productive.

SDC is therefore a reliability problem that the detect→drain→restart loop cannot solve on its own, because there is nothing to detect by conventional means. The answer is a separate detection program — fleet-wide hardware scanners run on idle cycles, in-production checkers that sample computations for correctness, and periodic re-validation against known-good references (Meta's Fleetscanner/Ripple/Hardware Sentinel family is the public reference design). When SDC is found, preserve suspect outputs and identify the last trusted validation boundary; quarantine the implicated node while hardware, software and data attribution determine its disposition, but the detection has to come from a dedicated program, not the recovery loop. The full SDC taxonomy, detection mechanisms, and fleet data are canonical in Chapter 14.3; the telemetry that surfaces it lives in Chapter 10.6.

Anti-patterns

The same reliability mistakes recur, because each comes from optimizing one number in isolation rather than goodput end-to-end:

  • Choosing topology from checkpointability. Checkpoint-and-resume lowers the consequence of some interruptions but does not prove that N/N+1 is sufficient or that 2N is waste. Compare the named facility states, interruption, post-event loading, recovery, independence, common modes, and contract with checkpointing and spares (Chapter 12.5).
  • Optimizing MTBF instead of MTTR. Pouring the reliability budget into hardware screening to push MTBF up a few percent, while restart still takes twenty minutes. At a fleet failing every few hours, compare the loss avoided by halving MTTR with that avoided by fewer interruptions, including lost progress and reserve. The larger return depends on the measured clocks and the intervention cost, rather than a universal preference for software.
  • An autonomous loop without lemon tracking or rate limits. Auto-rebooting and re-admitting nodes with no health history and no remediation rate limiter, so a single lemon flaps the job and a correlated firmware fault triggers a fleet-wide remediation storm.
  • Treating SDC as someone else's problem. Relying on the hang-and-restart loop to catch corruption it is structurally blind to, and discovering days of poisoned training only at the next eval. SDC needs a dedicated detection program (Chapter 14.3), not a louder watchdog.

Buy recovery time from the stage that dominates the measured critical path, with its spare, checkpoint and ownership dependencies inside the test. Faster storage earns nothing while allocation is blocked; more spares earn nothing when the checkpoint cannot resume valid work. Charge the remedy against 14.1, using 9.4 for checkpoint semantics.

This chapter is the operational hub of the reliability stack; the depth lives in its neighbors. The checkpoint anatomy and Young/Daly optimal-interval math are canonical in Chapter 9.4. The failure taxonomy (hard / transient / silent), the SDC detection program, and the empirical fleet AFR data are canonical in Chapter 14.3, with operational tuning — lemon-node ejection, elastic training in production, MTTR decomposition — in Chapter 14.4. The telemetry that feeds detection (DCGM/NVML, XID/SXID, goodput as the headline metric) is owned by Chapter 10.6. The availability-vs-goodput rethink that this chapter takes as its scoring lens is the subject of Chapter 12.2, quantified in Chapter 12.5 and contracted in Chapter 12.4. For inference fleets, multi-region failover is in Chapter 12.3 and the serving SLOs in Chapter 10.11. The redundancy-tier choices these reliability postures imply trace back to the workload archetypes of Chapter 1.1.
Cite this chapter
Fehn, J. (2026). Fleet Reliability, Fault Tolerance & Autonomous Recovery (Chapter 10.7). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-7-fleet-reliability-fault-tolerance-and-autonomous-recovery (accessed 2026-09-29).
@misc{aidc-10-7,
  author       = {Fehn, Jacob},
  title        = {Fleet Reliability, Fault Tolerance & Autonomous Recovery (Chapter 10.7)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-7-fleet-reliability-fault-tolerance-and-autonomous-recovery},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit