The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 14.3

In this chapter · 5 sections
Term help

Component Failure Modes, Failure Rates & Fleet Reliability Data

A large GPU fleet accumulates hard, transient and silent failures; Meta’s every-few-hours job interruptions illustrate one operating regime, so measure each component’s exposure and the whole-job trace before engineering detection, spares and recovery.

GOODPUTDENSITY-RAMPPOWER-BOUND

What you'll decide here

  1. Which failure taxonomy you instrument for — and specifically whether you fund a silent-data-corruption (SDC) detection program at all, because the failures you do not look for are the ones that quietly poison a training run for days.
  2. What annualized failure rate (AFR) you use per component — from unique failed FRUs and measured exposure — and whether Meta/SemiAnalysis evidence matches that cohort; spares also need replacement demand and lead time, while the SLA inherits a whole-service model.
  3. How long you burn in before accepting a node into production — the trade between schedule (every burn-in hour is deferred revenue) and infant mortality landing on the customer's training run.
  4. Where you draw the line between a node worth repairing and a 'lemon' worth ejecting — the ejection threshold that turns a long tail of repeat-offenders into reclaimed goodput.
  5. Which reliability evidence feeds each downstream model: unique failed FRUs over population-time for equipment AFR and spares, versus whole-job interruption events over runtime for checkpoint and goodput simulation; neither denominator may be transformed into the other.
Larger synchronous jobs expose more failure domains, but interruption cadence must be measured for the named fleet, job, and event definition — calibrate from telemetry rather than publishing a minute-level extrapolation.

This is the canonical home for one uncomfortable fact: a large AI training cluster must expect component interruptions during a run and recover from them. A traditional enterprise data center measures uptime in nines and counts annual outages on one hand. A 16,384-GPU training cluster experiences an unplanned interruption roughly every three hours — Meta's published Llama 3 405B run logged 419 of them over 54 days — and a well-run fleet still delivers over 90% effective training time through automation, not through preventing the failures (Meta, Llama 3 Herd, 2024). At this density and scale, the per-component physics guarantees components fail; the reliability problem is measuring the rate precisely enough to size spares, model availability, and detect the failures that hide.

This chapter establishes three things every other reliability chapter in the guide depends on. First, the failure taxonomy — hard, transient, and silent — which is the canonical fault vocabulary cross-referenced from the GPU operations view in Chapter 10.7, the redundancy view in Chapter 12.1, and the IST failure-demonstration view in Chapter 13.6. Second, the empirical fleet failure-rate data — the actual published numbers from at-scale operators, with their vintages and caveats, that you plug into an availability model rather than inventing. Third, the AFR modeling and burn-in discipline that turns raw component failure rates into a spares forecast and an acceptance gate. Every AFR in this chapter feeds the cluster availability and goodput roll-up in Chapter 12.5; the consolidated FMEA catalog these modes populate lives in Appendix F.

The failure taxonomy: hard, transient, silent

Every fault in a GPU fleet falls into one of three classes, and each demands completely different detection and recovery machinery; confusing them is a common operational error. The taxonomy organizes everything downstream: what you instrument, how fast you must react, and whether the failure is even visible at all.

Hard failures are the easy ones, paradoxically, because they announce themselves. A GPU throws an uncorrectable XID and falls off the PCIe bus, an optical transceiver goes dark, a CDU trips on low flow, a power supply faults. The component is unambiguously dead or unreachable; the job crashes or the node drops out; the telemetry screams. XID 79 ('GPU has fallen off the bus') is an archetypal hard-failure signal, but this guide does not assign it a portable fleet prevalence without a named equipment population, exposure window and study. These are expensive in lost goodput but cheap to detect: the recovery path is fail-fast, drain, restart from checkpoint, swap the FRU. Hard failures are where mature operators are already good, because the signal is loud.

Transient failures are the hard middle case. A correctable ECC error storm on HBM, a link that flaps rather than goes down, a thermal excursion that throttles a GPU for ninety seconds, an XID that self-clears on reset. The component is not dead — it works again after a power-cycle or a few minutes — but it degraded the run while it misbehaved, and it will very likely do so again. Transients are the raw material of lemon nodes: hardware that passes every point-in-time health check yet fails repeatedly under load. The decision a transient forces is not 'is it broken' but 'is it broken often enough to eject' — and getting that threshold wrong either keeps a repeat-offender poisoning runs or ejects healthy capacity. Correctable-ECC trends are one degradation signal to combine with XIDs, row-remap state, thermals, and load-correlated recurrence before deciding whether to eject hardware.

Silent failures are the dangerous class, because by definition nothing screams. Silent data corruption (SDC) is a computational error — a multiply that returns the wrong product, a memory read that flips a bit undetected by ECC — that produces no fault, no XID, no log line. The hardware reports success. The math is wrong. In training, an SDC silently corrupts gradients and weights; the loss curve drifts or diverges days later, and you cannot tell whether it is a bad hyperparameter, a data bug, or a single faulty multiplier in one of a hundred thousand chips. Meta’s 22 July 2025 reliability discussion reports about one SDC fault per thousand devices, without a stated exposure period. Its comparison with soft-error-related bit flips concerns a different error class, not a measured rise in one SDC rate. On a 100,000-device fleet, that statement cannot establish the fraction of defective machines or the probability of corruption during a multi-day run. Fund a deliberate detection program around the instructions, data patterns and operating conditions it covers; budget the GPU-hours for screening and the trusted progress lost before detection.

The three failure classes and their operational consequences
ClassSignatureDetection mechanismReaction timescalePrimary risk if missed
HardUncorrectable XID, device off bus, link down, hardware faultXID/SEL/syslog, DCGM, fabric BER alarms, CDU/PDU telemetrySeconds — fail-fast, drain, restart from checkpointLost wall-clock to last checkpoint; one node stalls the whole synchronous job
TransientCorrectable-ECC storm, link flap, thermal throttle, self-clearing XIDTrend analysis on correctable errors; repeat-offender counters; straggler detectionMinutes to hours — quarantine, observe, decide eject vs keepLemon node poisons run after run; a missed degradation signal before a hard failure
Silent (SDC)No signature — correct-looking but wrong computationDedicated SDC program: periodic test sweeps + in-workload checks + anomaly detectionDays — only surfaces as drifted/diverged training or wrong inference outputCorrupted weights, wasted compute, results you cannot trust; root-cause is days of detective work
Each class demands its own detection machinery, reacts on a different timescale, and is referenced as the canonical fault vocabulary from Chapters 10.7, 12.1, and 13.6.

Component failure modes: where the rate actually comes from

Fleet-level AFR is an aggregate that hides a strongly skewed distribution: a handful of components dominate the failure budget, and knowing which ones lets you target spares, burn-in, and detection where they pay off. The Meta Llama 3 root-cause breakdown is a useful public job-interruption dataset, but neither its counts nor its category mix is a portable equipment failure-rate distribution.

The GPU and its HBM dominate. In the Llama 3 paper's Table 5, 148 interruptions were attributed to faulty GPUs and 72 to HBM3 among 419 unplanned interruptions on one 16,384-H100, 54-day run. These are event counts, not unique failed devices or component exposure; the paper's printed percentages are internally inconsistent, so use the counts and preserve the named job boundary. This is not surprising once you see the physics: the accelerator package is the densest, hottest, highest-current component in the rack, and HBM stacks are the most thermally and mechanically stressed memory ever shipped at volume. Trend HBM health jointly with junction temperature, workload, ECC events, and remap history; as rack power rises, preserve thermal margin through the selected coolant-temperature, flow, and cold-plate envelope in Chapter 1.2.

Network and optics are the persistent long tail. Switches and cables accounted for 8.4% of Llama 3 interruptions, and link-flaps are as damaging as hard-down links because they corrupt collectives without obviously failing. At 800G XDR and the optics densities of a rail-optimized fabric, transceiver and cable failures scale with link count — a 100k-GPU cluster has millions of optical links, and even an excellent per-link AFR multiplies into a steady drip of fabric faults. The fabric is the failure domain that grows fastest as you scale out.

Infrastructure failures are rarer but higher-impact. Power and cooling faults are far less frequent than GPU faults per-event, but a single CDU trip or a PDU fault can take down an entire rack or pod at once — converting one component failure into dozens of simultaneous node losses. The Uptime data is stark: power is implicated in roughly 45% of impactful data-center outages (mostly UPS), and Uptime Institute's 2025 human-error survey findings use narrower denominators: 58% of human-error outages involved failure to follow an established procedure, while roughly 85% involved either that behavior or a flawed procedure. Neither figure is a share of all serious outages. The lesson for AI fleets is that the GPU dominates frequency while infrastructure dominates blast radius — and your FMEA in Appendix F must weight both.

Llama 3 405B interruption root-cause breakdown (16,384 H100s, 54 days)
Root causeShare of interruptionsClassSpares / detection implication
Faulty GPU (incl. XIDs)30.1% printed (148 of 419 events = 35.3%)Mostly hard, some transientLargest single spares driver; trend ECC and remap history; eject against calibrated thresholds
HBM3 memory17.2%Hard + thermal-drivenCoolant temperature directly modulates this term; bin GPUs with HBM history
GPU SRAM4.5%Hard/transientOften surfaces as correctable-error storms first
GPU processor4.1%HardBin GPUs with error history; size spares from unique failed FRUs over measured exposure, never from these interruption counts (148 GPU + 72 HBM3 of 419), whose printed percentages do not reconcile with them
Network switch / cable8.4%Hard + link-flap transientScales with link count; optics spares pool sized to fabric, not node, count
Software / other~12.9%TransientNot a spare; recovered by restart, masks some hardware root causes
The canonical public failure-mix dataset. 466 total interruptions, 419 unplanned, ~1 every 3 hours; >90% effective training time achieved with only 3 manual interventions. A single snapshot whose printed counts and percentages do not reconcile; do not generalize its failure mix to another fleet. Source: Meta, Llama 3 Herd of Models (2024).

The scale law: why MTBF collapses as the cluster grows

Job interruption cadence belongs to a named source population and event definition, not to GPU count alone. Meta measured 419 unplanned interruptions over 54 days while training Llama 3 405B on 16,384 H100s — about one every three hours for that run, including hardware and software causes. A separate SemiAnalysis reference reports roughly seven days MTBF for one 512-H100 cluster at a top-tier operator. These are useful anchors, but they are not points on one arithmetic curve: fleet maturity, topology, job membership, correlated faults, software, detection policy, and the interruption denominator differ.

The stable scale lesson is directional: a larger synchronous job participates in more failure domains, so recovery and GOODPUT matter more as the job grows. Fit the effective job-level failure distribution from fleet telemetry, then compute checkpoint cadence from that measured distribution and the actual save/restart cost; do not infer a universal two-minute interval or a minute-level failure forecast from accelerator count. The checkpoint math is canonical in Chapter 9.4; operational tuning is in Chapter 14.4.

1 every ~3 hr
unplanned interruption rate, Llama 3 405B (419 over 54 days, 16,384 H100s); >90% effective training time with automation
148 GPU + 72 HBM interruptions
interruption attribution in one 16,384-H100, 54-day Llama 3 405B run: 148 faulty-GPU and 72 HBM3 events of 419 — event counts, not unique failed devices
Scope & caveats

Table 5's printed percentages are internally inconsistent; retain the event counts and the named 16,384-H100, 54-day job boundary. Counts are not unique failed FRUs or equipment AFR.

Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.

148 GPU + 72 HBM interruptions
observed interruption attribution in one 16,384-H100, 54-day Llama 3 run: 148 GPU + 72 HBM events; not equipment AFR
Scope & caveats

Job-interruption event counts for one named run. The source does not establish unique failed devices or equipment population-time exposure, so these counts must not be annualized into component AFR, fleet lambda, cumulative equipment risk or spares demand.

~1 fault per 1,000 devices; exposure unstated
Meta reported about one SDC fault per 1,000 devices; exposure period unstated (22 July 2025)
Scope & caveats

Historical statement in Meta’s AI-hardware reliability discussion; no time denominator is supplied. Not a measured annual incidence, per-job cadence or proof that 0.1% of GPUs currently harbor a defect.

45–60 days
Meta CPU-fleet full sweep cadence, July 2025 report; not GPU/HBM coverage
Scope & caveats

CPU-fleet detection example. Coverage cadence is not detection latency for every defect or GPU/HBM diagnostic coverage. Hardware Sentinel coverage ratios are a distinct result.

72–168 hrguidance
2025 practitioner example: 72–168 hr GPU node soak; not a universal acceptance bound
Scope & caveats

Together AI and ClusterMAX described a 72–168 hr range in 2025 practitioner guidance. The project/OEM/contract test plan sets duration, intensity, statistical stopping rule, pass/re-soak gate, and vendor disposition.

~7 days / one 512-H100 cluster
reported MTBF for one 512-H100 cluster at a top-tier operator
Scope & caveats

SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.

SDC detection programs: chasing the failure with no signature

Because SDC by definition leaves no log line, detecting it is an active program, not a passive alarm — and the state of the art is a layered defense, with each layer trading coverage against the GPU-hours it steals from production. Meta's published CPU-fleet stack illustrates that tradeoff; it does not establish GPU-fleet coverage.

Fleetscanner is the offline sweep: dedicated CPU test patterns scheduled across Meta’s fleet on a reported 45–60-day CPU sweep cadence. It is the most thorough layer (over three years it reached ~93% coverage for a major defect family, with ~23% unique coverage no other method caught) but the most expensive, because the CPU under test is unavailable for its production workload. Ripple co-locates with live workloads, slipping millisecond-to-second test bursts into the gaps between real work, so it achieves fleet-wide coverage in days rather than weeks at near-zero opportunity cost — at lower per-pass depth. Hardware Sentinel is the newest layer: it watches application exceptions in kernel space and infers core-level SDC without allocating any test time at all, reporting effective CPU defect coverage roughly 1.74x over Fleetscanner and 1.92x over Ripple in the studied population (ASPLOS 2025); neither ratio is GPU/HBM coverage. The architectural lesson is that no single method suffices — you layer a deep-but-slow sweep, a fast-but-shallow in-workload probe, and an application-exception inference layer with monitoring and validation costs, and the union catches what any one misses.

For training specifically, the framework-level defenses matter as much as the fleet-level ones: redundant computation on a sample of operations, gradient/activation checksums, and divergence monitors that flag when a replica's numerics drift from its peers. These can catch covered SDC paths during computation; they do not detect every wrong result, and the CPU sweep’s 45-day clock is not a GPU-diagnostic delay estimate. The decision here mirrors the funding fork above — every layer you add costs GPU-hours or engineering, and the right depth is set by how catastrophic a silently-corrupted run is for your business. A frontier lab burning months of compute on one run has a strong reason to fund layered detection; a batch-inference shop can buy less only if output checks and replay contain silent errors before they escape. Neither workload can infer GPU coverage from Meta’s CPU results.

Deep dive: from component AFR to a spares forecast and an availability number

The practical payoff of measuring failure is a defensible spares forecast and availability input, but each rate needs a matching denominator. Estimate equipment AFR from unique failed FRUs over equipment population-time, segmented by generation, age, duty and failure definition. Separately measure job interruptions over job runtime for checkpoint and goodput modeling. Meta's 148 GPU-attributed and 72 HBM-attributed interruptions among 419 events on one 16,384-H100, 54-day Llama 3 run describe that job's interruption mix; repeated events, restored devices and exposure are not resolved, so those counts cannot be annualized into equipment AFR, fleet swaps or a 100,000-GPU spares premise.

Feed the population-matched equipment estimates into the availability and replenishment models, and the measured whole-job process into goodput simulation. The full sparing model, RMA logistics, and repair-vs-replace-vs-harvest economics are in Chapter 14.6; the roll-up methods are in Chapter 12.5.

Burn-in: paying for infant mortality up front

Component failure rates are not constant over life — they follow the classic bathtub: a high infant-mortality phase early, a low flat useful-life phase, and a rising wear-out phase late. Burn-in is the deliberate decision to pull infant mortality forward into a controlled acceptance window so it lands on a test harness instead of a customer's training run. It is a direct schedule-versus-reliability trade, settled at go-live.

The campaign is selected from the signed OEM/project/contract plan: declared stress profile and cycles, fleet exposure, measured discoveries by failure mode and node-hours, pass/re-soak dispositions, and a pre-agreed statistical stopping rule. Together AI and ClusterMAX described 72–168 hours in 2025 practitioner guidance; that is an example, not a universal bound. DCGM diagnostics can be components of the evidence package, but no run level alone is a universal acceptance gate. A separate October 2024 SemiAnalysis playbook recommends at least 3–4 weeks of factory high-temperature burn-in before deployment; it does not report a universal early-failure settling period. A burn-in campaign's claimed early-failure removal is only usable with its tested population, protocol, warranty interval and source locator; without those, treat the campaign as an acceptance gate, not a quantified reliability guarantee.

The burn-in duration fork: schedule vs. reliability
Burn-in postureWindowWhat it buysWhat it costsBest fit
Smoke test only<24 hr exampleFastest releaseLittle evidence about time-dependent or intermittent defectsUse only where the signed risk basis explicitly accepts that residual exposure
Dated practitioner example72–168 hr (2025 example)A repeatable campaign window when its stress profile and stopping evidence are declared3–7 days of deferred revenue per cohortNot a default; adopt only when the project/OEM/contract plan justifies it
Evidence-extended campaignContinue to contracted stopping ruleMore exposure and confidence where discoveries or uncertainty remainMore deferred revenue and test wearAny fleet whose signed evidence rule has not yet cleared
These rows are planning postures, not duration prescriptions. The signed plan declares stress intensity, fleet exposure, discovery confidence, pass/re-soak rules, and the statistical stop condition.

Burn-in does not end at acceptance — it transitions into a steady-state cadence. One possible day-2 schedule pairs a weekly deep node-health pass (dcgmi diag -r 3 plus NCCL on drained GPUs) with the continuous straggler/lemon detection that watches for the repeat-offenders burn-in could not catch; the actual cadence follows supported diagnostics, fleet evidence and the maintenance budget. A site may investigate a node running ~15% below a golden-reference benchmark as an illustrative trigger, after matching workload, power cap, temperature and software; automatic quarantine needs a validated threshold and false-positive budget, and the lemon-ejection decision — proven to cut 512+-GPU job failure rates from ~14% to ~4% and lift completion ~30% (Meta lemon-node studies, 2024) — is where transient-failure data becomes reclaimed goodput. Burn-in front-loads the cost of infant mortality; lemon ejection back-stops the transients that slip through. Both are detection programs paid for in GPU-hours, and both need the same accounting: compare reclaimed goodput with diagnostic GPU-hours, false quarantines and engineering cost in the 14.1 ledger.

Deep dive: why your fleet's numbers will (and should) differ from Meta's

The Llama 3 dataset is the most-cited reliability data in the field precisely because so little else is public — but treating it as a universal constant is a mistake. It is a single snapshot, on H100s, on Meta's specific facility, cooling, firmware, and software stack, in 2024. Four things move your numbers off it. Silicon generation: Blackwell-class GB200/GB300 racks at ~132–142 kW change the thermal and current stress profile entirely, and their burn-in AFR is still being established across the fleet — early NVL72 bring-up surfaced novel reliability issues that did not exist on H100 — by 2026 SemiAnalysis had pinned the dominant one specifically on the compute tray's flyover/ACC cables (not just the copper backplane), calling cable terminations the #1 failure point of GB200/GB300 assembly; Meta's custom "Ariel" GB200 NVL72 amplified it into cross-rack NVLink signal-integrity failures and is reverting to a standard Oberon design. NVIDIA's answer in Vera Rubin NVL72 is a cableless compute tray (blind-mate board-to-board connectors, assembly time cut from ~2 hours to ~5 minutes) — a reliability bet as much as an assembly-speed one. Cooling discipline: set intervention thresholds from the fleet's joint junction-temperature, workload, ECC, and remap history. Operational maturity: a fresh cluster in burn-in and a two-year-old fleet sit on opposite ends of the bathtub curve, so a blended fleet AFR depends on your age mix. Software stack: Llama 3's ~12.9% software share is highly stack-dependent and not portable at all.

The conclusion is operational: borrow only population- and event-definition-matched evidence to bootstrap a design-time model, then replace it with measured unique-FRU exposures and whole-job event distributions from the named fleet. Llama 3 job-interruption counts may inform an interruption-mix scenario; they do not bootstrap equipment AFR. The DCIM and observability stack of Chapter 14.2 exists in large part to produce your AFR, not someone else's — and the availability model in Chapter 12.5 is only as good as the fleet-measured failure rate you feed it.

This chapter is the canonical home for the hard/transient/silent taxonomy and the empirical failure-rate data; the modes it catalogs populate the consolidated FMEA in Appendix F. The same fault vocabulary is used operationally in Chapter 10.7 (fleet fault tolerance & autonomous recovery), in the redundancy and fault-domain engineering of Chapter 12.1, and demonstrated under load in the Level 5 IST failure-mode work of Chapter 13.6. The AFRs derived here feed the quantitative availability and goodput roll-up in Chapter 12.5. The checkpoint-interval math that turns the scale law into a recovery strategy is canonical in Chapter 9.4 and operationally tuned in Chapter 14.4. The spares and RMA logistics that consume these failure rates are in Chapter 14.6; the telemetry that measures them in Chapter 14.2; and the density ramp that keeps invalidating the data in Chapter 1.2.
Cite this chapter
Fehn, J. (2026). Component Failure Modes, Failure Rates & Fleet Reliability Data (Chapter 14.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-3-component-failure-modes-failure-rates-and-fleet-reliability-data (accessed 2026-09-29).
@misc{aidc-14-3,
  author       = {Fehn, Jacob},
  title        = {Component Failure Modes, Failure Rates & Fleet Reliability Data (Chapter 14.3)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-3-component-failure-modes-failure-rates-and-fleet-reliability-data},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit