Chapter 14.4
In this chapter · 6 sections
Reliability Engineering for Training (Operational)
Where a frontier training job is interrupted every few hours, as in Meta’s named run, operations turns on detecting, isolating and recovering faster than failures accumulate — every recovery minute costs goodput on the declared job ledger.
What you'll decide here
- Your checkpoint tier mix and operational cadence — how much you spend on RAM/peer/persistent tiers to drop mean-time-to-recovery, recalculating the Young/Daly interval when measured checkpoint cost or interruption/recovery inputs change (the math lives in Chapter 9.4; here you move ETTR).
- Where the fault-detection boundary sits — what you catch with synchronous health checks at job launch versus passive monitoring mid-run versus a periodic offline node-sweep — and how aggressively you eject suspected lemon nodes before they have proven themselves bad.
- Whether you run rigid (fail-stop-and-restart) or elastic/redundant training — buying continuity through spare capacity and reconfiguration against the goodput tax of running with hot spares idle.
- Who owns the detect-to-recover loop and how automated it is — the facility-ops/ML-platform boundary, and how far you push autonomous remediation before a human is in the loop.
- Which MTTR component you attack next — detection latency, isolation/scheduling delay, reconfiguration, or reload-and-replay — because they have wildly different costs to shave and the binding one moves with cluster size.
A frontier training job is a single synchronous computation spread across tens of thousands of accelerators, and it advances at the speed of its slowest healthy participant. When any one of those participants dies, the whole job stops. Training reliability is unlike any other operational discipline in the building: there is no graceful degradation, no shedding of load, no failing-over of a request. Meta's 16,384-GPU Llama 3 run recorded 419 unplanned interruptions over 54 days — about one every three hours for that named run. Do not turn that observation or RSC-2 component rates into a 131,072-GPU minute-level forecast without explicit population, job-membership, independence/common-mode, software, and event-definition assumptions. At that scale the question is no longer whether the run will be interrupted but how much wallclock you burn each time it is — and that number, summed over a multi-week run, is the difference between a 70% and a 95% effective-training-time job.
This chapter is the operational counterpart to the design-time reliability work in Part 12 and the failure-data catalog in Chapter 14.3. It is not where checkpoint interval math lives — the Young/Daly derivation is canonical in Chapter 9.4, and useful-output accounting is canonical in Chapter 14.1, with service boundaries in Chapter 12.2. Here we cover the four operational levers an ops team actually turns on day-2: the checkpoint tiering that sets how fast you recover, the detection and isolation that decides which node is at fault and ejects it, the elastic-vs-rigid recovery posture, and the MTTR decomposition that tells you which of these to spend the next dollar on.
The operational target: ETTR and service continuity
The facility world measures itself in availability nines — the legacy Tier III/IV mappings still quoted in every sales deck (Uptime itself has disavowed them). For a synchronous training job, facility availability constrains continuity but does not tell whether its specific 16,384 GPUs made forward progress on the same step. The operational metric is Effective Training Time Ratio (ETTR): productive runtime divided by the job's available wallclock — the same denominator as Chapter 14.1, scheduled or eligible-queue time, with restart overhead included; Google application goodput starts at the application boundary and must not silently inherit scheduler queue time. Equivalently, the industry speaks of goodput (forward progress) net of badput (restarts, re-computation since the last checkpoint, idle time waiting for a replacement node, slow stragglers). Treat 0.90 and 0.96 ETTR as illustrative sensitivity points; well-run RSC-1 jobs at 2,048–4,096 GPUs exceed 0.90 even in a congested shared cluster, while an unoptimized 16,000-GPU job projects to only ~0.70 with naive checkpointing — a 23-point goodput gap that frequent checkpointing and fast restart close to ~0.93.
The reason this matters in 2026 is economic: the cluster is power-bound, and the GPUs are depreciating against a contested 2–3-year bear-case economic clock whether or not they are doing useful work. The April 2026 ClusterMAX model’s ~6–21% goodput expense includes assumed recovery and fault-tolerance overhead; measure the job’s loss account before pricing it. On a fleet earning on the order of $12–13B per GW per year (a contested, single-source figure), a 5-point ETTR improvement is the difference between hitting and missing the run's compute budget, but it creates usable compute capacity, not automatically recovered capital or cash. Account for lost progress in Chapter 14.1; value recovered capacity against paid demand and contribution in Chapter 1.3/1.8.
Checkpoint tiering in operation
The optimal checkpoint interval is a solved problem — Young's first-order approximation uses checkpoint cost and MTBF, while Daly's higher-order model adds recovery, restart, and downtime terms, and that derivation is canonical in Chapter 9.4. What the interval math takes as an input, and what operations actually controls, is the cost of a checkpoint and the cost of a restart. Drive those down and the optimal interval shrinks, the work-at-risk between checkpoints shrinks, and ETTR rises — without buying a single additional GPU. The lever for that is tiering: keeping checkpoints at multiple storage levels so that the common-case recovery never has to touch the slowest, most durable tier.
For the stated mixed-precision model-and-optimizer layout of 14 bytes per parameter, a 405B-parameter checkpoint is ~5.7 TB before metadata and redundancy; other state layouts change that payload. Writing that to a persistent parallel file system on every interval is bandwidth-prohibitive and stalls the job; reloading it across the fabric on every failure is the dominant restart cost. Tiering solves both. The canonical three-tier scheme keeps the freshest checkpoint in node-local RAM/host memory (potentially faster reload when the job restarts on the same hardware; measure complete restart latency, not memory access time), a redundant copy on a peer node or adjacent slice (survives a single-node loss without touching durable storage), and a periodic copy to durable object/parallel storage (survives a whole-job or facility event). Google reported a 6.59% goodput uplift on a 35K-chip TPU v5p workload, checkpoint-save latency under five minutes, and restore under one minute across thousands of nodes; AWS separately reported checkpointless peer recovery cutting 15–30 minutes to under two minutes on 2,304 H100 GPUs.
| Tier | Medium | Recovery latency | Survives | Operational cost |
|---|---|---|---|---|
| In-memory / host RAM | Node-local DRAM (and HBM staging) | Seconds — no network reload | Transient process/GPU error, same-node restart | DRAM capacity reserved off the training footprint; lost if node dies |
| Peer / in-cluster | Replica on a neighbor node or adjacent slice | Tens of seconds — intra-fabric copy | Single-node loss without a durable reload | Extra fabric traffic + 2x checkpoint memory; erasure-coding reduces it |
| Durable / persistent | Parallel FS or object store (GPUDirect Storage) | Minutes — full reload across the fabric | Whole-job, rack, or facility event | Storage bandwidth; the stall the async drain is hiding |
| Async drain (overlay) | Background copy from RAM to durable | N/A — hides write cost | Makes durable-tier cadence affordable | Engineering complexity; <10% compute overlap target |
Two operational refinements sit on top of the tiers. Asynchronous checkpointing snapshots state to host memory in a brief synchronous pause, then drains it to durable storage in the background while training continues — the target is keeping the visible stall under ~10% of the checkpoint window, which is what makes a sub-5-minute durable cadence affordable at all. Sharded / distributed checkpoint formats let each rank write its own shard in parallel, so checkpoint wall-time scales with per-node bandwidth rather than aggregate model size — without it, the checkpoint itself becomes the straggler. The operational fork is how much memory and fabric you are willing to spend on the upper tiers: a job that checkpoints only to durable storage is simple and cheap to run but recovers slowly; a fully-tiered job recovers in seconds but reserves host RAM and burns fabric bandwidth on replica writes. The right answer is set by where your MTTR decomposition (below) says the time is going.
Deep dive: why the restart cost, not the checkpoint cost, is the operational lever
Operators new to training reliability instinctively optimize the write side — faster checkpoints, higher storage bandwidth. But checkpoint cadence depends on measured whole-job interruptions and both checkpoint and restart costs, so write bandwidth is only one lever. The bigger lever is the restart cost, because it is paid in full on every single failure, and at scale failures are frequent. Restart cost has four serial components: detect the failure, isolate and reschedule onto healthy hardware, reconfigure the parallel topology, and reload-plus-replay from the last checkpoint. Tiering attacks the reload term (load from RAM, not from the parallel FS); fast detection attacks the first term; hot spares and elastic reconfiguration attack the middle two.
The target is fleet-specific. Measure the effective interruption distribution for the named job, then solve for checkpoint cadence and restart overhead using the actual ETTR objective and save/restart costs. Hyper-checkpointing and in-memory recovery can keep the common-case path off durable storage, but a universal two-minute target does not follow from accelerator count or an RSC-2-like component rate alone. The operational implication: if your detect-to-recover loop is dominated by the durable reload, you are tiering wrong, and no amount of faster checkpoint writes will save you. → Chapter 9.4 for the interval derivation.
Fault detection, isolation, and lemon-node ejection
A restart is only as good as the decision about where to restart. Resume the job on the same flaky node and it fails again within the hour; the run thrashes, ETTR collapses, and on-call burns out. So the second operational pillar is detection and isolation: catching that a failure occurred, attributing it to the right component, and removing the bad hardware from the schedulable pool before it poisons the next attempt. The hardest cases are not the hard failures (a GPU that falls off the bus is loud and easy) but the silent and gray ones — silent data corruption that produces wrong gradients with no error, and stragglers that are technically alive but running 20% slow and dragging the whole synchronous job down to their pace. The failure taxonomy (hard / transient / silent) is canonical in Chapter 14.3; here the concern is the operational detection program that sits on top of it.
Detection runs at three boundaries. Synchronous health checks at job launch (and after every restart) sweep the allocated nodes for known-bad signatures before training starts — cheap, but only catch what they test for and add latency to every recovery. Passive in-run monitoring watches per-step timing, collective-op latency, ECC counters, and thermal/throttle telemetry to flag stragglers and degrading nodes mid-run — catches the gray failures the launch check misses, but risks false positives that eject healthy nodes. A periodic offline node-sweep qualifies idle nodes against a benchmark to find slow/SDC-prone hardware before it is ever scheduled — the most thorough, but consumes capacity that could be training. Mature operators run all three; the question is the weighting.
Isolation feeds automated remediation: once a node is flagged, the workflow drains the job off it, attempts an automated recovery ladder (reset the GPU, reload the driver, power-cycle the node, reflash firmware), re-qualifies it against the health suite, and either returns it to the pool or opens an RMA and pulls a spare. The degree of automation here is the facility-ops/ML-platform-ops boundary made concrete: the platform owns the detect-flag-eject loop and the job-level recovery, while facility ops owns the physical swap and the RMA logistics. Drawing that boundary cleanly — and deciding how far the autonomous ladder runs before a human is paged — is an organizational decision treated in Chapter 14.11; the autonomous-recovery mechanics overlap with Chapter 10.7.
Scope & caveats
Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.
The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.
Scope & caveats
Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.
Scope & caveats
Named RSC deployment and event definition; a useful external comparison, not a required calibration target for other fleets.
Scope & caveats
Google reports the figure as a 6.59% increase in ML Goodput on one 35K-chip TPU v5p workload; that wording does not establish 6.59 percentage points absolute. The save and restore latencies are checkpoint-path measurements, not an end-to-end detection-to-resumed-training MTTR for an arbitrary GPU fleet.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Elastic and redundant training
The default recovery posture is rigid: the job fails, stops entirely, the scheduler re-allocates a full healthy set of nodes, and the run reloads the last checkpoint and replays. Simple, robust, and the right choice for most jobs — but it pays the full detect-isolate-reschedule-reload cost on every failure, and at six-figure GPU counts the reschedule term alone (waiting for the scheduler to find and provision a clean replacement set) can dominate. The two operational alternatives buy faster recovery by spending capacity.
Elastic training lets the job continue on the surviving nodes at reduced width — drop the failed node's data-parallel replica, reshard, and keep stepping — then re-expand when a replacement is qualified. It eliminates the stop-and-reschedule stall for the common single-node case, at the cost of running slower (and at slightly different effective batch size) until the node is back. Redundant / spare-pool training keeps hot spares qualified and idle so a failed node is swapped in seconds rather than waiting on the scheduler or an RMA — recovery approaches the reload time alone, but the spares are a standing goodput tax (idle GPUs that depreciate without computing). The fork is a direct trade between recovery speed and reserved capacity, and it interacts with the elasticity of the parallelism scheme — nonuniform / elastic tensor-parallelism reduces the failure-amplification penalty of a node loss, which is what makes degrade-and-continue viable at all.
| Posture | On a node failure | Recovery latency | Standing cost | Best fit |
|---|---|---|---|---|
| Rigid (fail-stop) | Whole job stops, full re-allocation, reload+replay | Full detect+reschedule+reload | None — simplest | Smaller jobs; abundant clean capacity; simple ops |
| Elastic (degrade-and-continue) | Reshard onto survivors, keep stepping at reduced width | Near-zero stall; runs slower until refilled | Throughput dip while degraded | Large jobs with elastic parallelism; staleness-tolerant |
| Redundant (hot spares) | Swap a pre-qualified spare in seconds | ≈ reload time only | Idle spare pool depreciating | Frontier runs where every restart is costly |
| Layered (elastic + spares) | Continue degraded, then swap spare, then re-expand | Lowest end-to-end | Spare pool + elasticity engineering | Six-figure-GPU jobs at best-in-class ETTR |
| Checkpointless / in-job peer recovery (HyperPod, TorchFT, NVRx) | Rebuild failed replica's state from peers/optimizer redundancy — no durable-storage reload | ≈ detect+reshard; AWS: 15–30 min → <2 min (90 s at 2,304 GPU) | Memory for redundant replicas; framework coupling | Large runs where storage-reload dominates MTTR (up to ~95% goodput, vendor-measured) |
MTTR decomposition: where the time actually goes
Everything above is in service of one number — mean-time-to-recovery — and you cannot improve it without decomposing it. Record these four phases on one incident timeline; sum only non-overlapping critical-path durations, each with its own owner, its own cost to shave, and its own scaling behavior:
- Detection latency — wallclock from the failure to the system knowing it failed. Dominated by how fast health monitoring fires; gray/silent failures (a slow straggler, an SDC) can take many steps to surface and are the worst offenders. Shave it with passive in-run monitoring and tighter step-time anomaly thresholds.
- Isolation & scheduling — attributing the fault to a component, ejecting it, and provisioning a healthy replacement set. This term grows with cluster size (more nodes to schedule around) and is the one hot spares and elastic continuation directly attack.
- Reconfiguration — rebuilding the parallel topology, re-establishing collectives, and re-warming the fabric on the new node set. Roughly fixed per restart; elastic/nonuniform parallelism reduces it by avoiding a full topology rebuild.
- Reload & replay — loading the last checkpoint and recomputing the work since it. This is the term tiering attacks (load from RAM, not durable storage) and the term Young/Daly's interval bounds (less work-at-risk between checkpoints).
Instrument all four and attack the binding one — because which term dominates shifts with scale. On a few-thousand-GPU job the reload-and-replay term usually dominates, so tiering and interval tuning pay off most. On a six-figure-GPU job the isolation-and-scheduling term grows until it rivals reload, which is why hot spares and lemon-node pre-ejection (keeping a clean pool so scheduling is instant) are where the next dollar goes. Optimizing the wrong term is the classic waste: buying faster checkpoint storage when your time is actually going to a slow scheduler.
Deep dive: the detect-to-recover loop as a closed control system
The most useful way to think about operational training reliability is as a closed-loop controller running continuously over the cluster, not a sequence of incident responses. The loop has a sensing stage (health checks, step-time telemetry, ECC/throttle counters, collective-latency probes), a decision stage (is this a transient to retry, a node to eject, or a straggler to fence?), an actuation stage (the automated remediation ladder and the scheduler), and a learning stage (every event feeds the lemon-node classifier and the fleet failure-rate model). The loop's bandwidth, how fast it goes from a degraded signal to a clean recovered job, is exactly the MTTR you are trying to minimize. Its precision — how often it ejects a healthy node or fails to catch a sick one — sets your false-eject capacity loss and your missed-failure replays.
Two design principles fall out of the control framing. First, detection must lead failure where possible: passive monitoring that flags a degrading node before it hard-fails turns an unplanned restart into a scheduled drain, collapsing the detection-and-scheduling terms. Second, the loop must learn: a lemon classifier that updates from each ejection, and a failure-rate model that updates from each swap, is what lets the threshold tuning above stay calibrated as the fleet ages through its bathtub curve. This is also where the operational twin closes the loop with as-built reliability data — the same feedback discipline as the design-validation twin in Chapter 2.7, applied to the running job rather than the building.
Operational anti-patterns
The recurring failures in training-reliability operations come from optimizing one term in isolation or from treating a frequent-failure regime as if it were a rare-failure one. Three are worth naming:
- Restarting on the same flaky node. No lemon-node ejection, so a degrading GPU re-kills the job every recovery cycle. The interruption count looks like bad luck; it is actually one bad node poisoning every attempt. The fix is cheap (eject before re-scheduling) and the payoff is large (14%→4% large-job failure).
- Tiering only to durable storage. In the cited AWS scenario, recovery without its faster peer path pays a 15–30 minute fabric reload; RAM or peer tiers can reduce that delay, but the installed storage layout determines the actual cost. At a few thousand GPUs this is survivable; at six-figure GPU counts, test whether reload, scheduling or reconfiguration binds before claiming 0.9 ETTR is out of reach. A memory/peer tier earns its cost when its measured restoration time and fault coverage improve the end-to-end ledger.
- Choosing facility topology from checkpointability. Checkpointing can reduce the lost work from an interruption but does not prove that a path or hall outage meets the recovery SLO. Compare facility states and common modes with faster recovery, hot spares, and lemon ejection in the Chapter 12.5 model.
Cite this chapter
Fehn, J. (2026). Reliability Engineering for Training (Operational) (Chapter 14.4). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-4-reliability-engineering-for-training-operational (accessed 2026-09-29).
@misc{aidc-14-4,
author = {Fehn, Jacob},
title = {Reliability Engineering for Training (Operational) (Chapter 14.4)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-4-reliability-engineering-for-training-operational},
note = {Accessed 2026-09-29}
}