The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 14.1

In this chapter · 7 sections
Term help

Operational KPIs, Goodput & the Reliability Economics of AI Factories

Goodput — accelerator-hours doing useful work on the critical path — governs an AI factory's day-2 economics; justify each reliability dollar by the useful output it buys while meeting the contracted service-availability obligation.

GOODPUTPOWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. Which top-line operating metrics govern the fleet — contracted service availability and ML goodput / ETTR, measured separately — because their boundaries re-weight every reliability investment downstream.
  2. What goodput target you contract and design for — and therefore how much you spend on fast checkpointing, hot spares, health-checking, and silent-corruption detection to close the illustrative 90%-versus-96% effective-training-time gap.
  3. How you set the blast-radius policy: how large a synchronous failure domain you tolerate before splitting it, given that NVL36x2 packaging spreads one 72-GPU NVLink domain across two racks rather than halving it, that one failed GPU in a tightly-coupled job idles the whole job and one bad node can cascade into preemptions across the cluster.
  4. Which reliability spend — recovery speed, lemon-node ejection, SDC scanning, 2N power or fault-tolerant cooling — preserves the most useful output while meeting the named maintenance/fault states and contract, including the cost of reserved capacity.
  5. Which operations scorecard the board and the customer see — the KPI set, its denominators, and the SLA definitions — because an un-named denominator (uptime of what, measured how) is where day-2 disputes and stranded-asset surprises hide.
Useful output (goodput) can peak below maximum admitted load when contention, restarts or missed SLOs consume the extra throughput; measure that knee for the workload before choosing utilization.

Parts 1 through 13 of this guide build the machine. This part operates it — and operating it is where most of the lifetime money is won or lost, because the asset depreciates whether or not it is producing. The reflex inherited from twenty years of enterprise data-center operations is to manage the facility against availability: nines of uptime, hours of downtime per year, Tier classes. That reflex measures continuity at its declared service boundary; it does not measure how much paid compute reaches the customer. An AI factory is a depreciating compute asset whose return is set by how many accelerator-hours land on the revenue-bearing critical path. A hall can report high facility availability and still throw away a fifth of its compute to stragglers, restarts, silent corruption, and idle hot spares — and the income statement will not care that the UPS never dropped.

This chapter is the framework for Part 14: it pairs contracted availability with goodput, defines the operating ledger (ETTR, ML goodput, MFU/MBU) at engineering depth, characterizes the AI failure environment that makes day-2 different from traditional IT, sets out the blast-radius problem that turns one component fault into a cluster-wide stall, and assembles the operations scorecard that the rest of Part 14 instruments. The service-boundary and continuity tradeoffs live in Chapter 12.2; this chapter owns the useful-output ledger — how you measure it, what it costs to move it, and what the scorecard looks like once the building is live.

Availability and goodput: choose the service boundary

The first day-2 decision is which number sits at the top of your operations dashboard, because it silently re-weights every reliability investment below it. Availability asks what fraction of the contracted time or requests met the named service conditions. Facility-power availability counts energized time; reachability counts reachable IT; correctness-aware service availability can also reject corrupt responses. Those boundaries require separate event ledgers, and Uptime Tier topology predicts none of their percentages. Goodput asks a different and harder question: what fraction of the accelerator-hours you paid for actually advanced a job on the critical path? The two diverge sharply for AI, and that divergence is what Part 14 instruments.

Consider a synchronous pre-training run on 16,384 GPUs. Suppose the facility never loses power during the run; that establishes continuity over this window, not an annual five-nines result. But the run is interrupted roughly every three hours at that scale — 419 unplanned interruptions over 54 days, from GPU, network, software and other causes (Meta's Llama 3 405B snapshot) — and each interruption idles all 16,384 accelerators until the job restarts from its last checkpoint. The 'available' facility is hemorrhaging goodput through a mechanism the availability metric cannot see, because the failure domain is the job, not the component. This is why the AI-cluster reliability model must include both facility states and workload recovery. Checkpointing, hot spares, and recovery automation may beat an additional facility path on marginal goodput for a named failure environment, but checkpointability does not prove that 2N is misallocated. Compare the defined maintenance and fault states, interruption, post-event capacity, independence, common modes, recovery, and contract before choosing the next investment.

The useful-output ledger: ETTR, ML goodput, MFU/MBU

'Goodput' is not one interchangeable number: ETTR, application goodput and MFU expose different losses, and conflating them can make a wasteful fleet look efficient. Read each numerator against its own denominator; multiply ratios only after proving that they are conditional factors on one event ledger and one time boundary, or the same lost hour is counted twice.

ETTR (Effective Training Time Ratio) measures productive runtime divided by the available wallclock time of a job run: time scheduled or eligible to be scheduled, including the eligible queue, ranging from 0 to 1, accounting for queueing delay, restart overhead, and re-computation of lost progress (Meta, Revisiting Reliability in Large-Scale ML Research Clusters, 2024–25). ETTR is model-agnostic — it does not care what fraction of a GPU's FLOPs a kernel extracts — which is exactly why it is the right SLA target between an operator and a training customer. Meta's largest reported jobs (>1024 GPUs) sustained average ETTR above 0.9 with one-hour checkpoint intervals on its named shared research clusters. A different fleet must derive cadence and restart targets from its own effective interruption distribution, checkpoint cost, recovery path, and ETTR objective; accelerator count alone does not produce a portable two-minute target.

ML goodput (Google's formulation) is the fraction of application elapsed time spent on preserved training progress, with all non-productive time accounted as 'badput': program startup, data-loading stalls, checkpoint writes that do not overlap compute, failed steps, wasted progress since the last checkpoint, and scheduling gaps. Google's ML Goodput library also measures an individual application's preserved training progress; its elapsed clock starts with the application, so it does not automatically include an earlier eligible queue. The operations scorecard rolls up compatible raw totals, weighted by allocated accelerator count when job widths differ. Google's August 26, 2026 documentation names this application boundary. Use 90% versus 96% as stipulated September 2026 sensitivity endpoints. Holding workload, peak-normalized productivity and the installed 1 GW IT capacity fixed, six points represent 60 MW-equivalent of useful-output capacity. The plant still draws its workload-dependent power: this is neither 60 MW saved electrically nor automatically sellable capacity. Chapter 1.3 maps qualified output to demand and revenue.

MFU (Model FLOPs Utilization) and MBU (Model Bandwidth Utilization) are the innermost layer — they measure efficiency while a job is running, independent of failures. MFU is the fraction of the hardware's peak FLOPs the model actually realizes (compute-bound training; Meta’s Llama 3 405B Table 4 reports 41% BF16 MFU on 16,384 H100s at 8K sequence in July 2024; published GB300 NVL72 measurements exist for named benchmark configurations — Lambda plots roughly 52.7–59.6% training Llama 3.1 405B at 8K sequence length — but no comparable fleet-wide operating band). MBU is the analogous ratio for memory bandwidth (the binding constraint on autoregressive inference decode). A job can have a 0.95 ETTR and a 35% MFU — perfectly reliable, half-idle on the math — which is why the three metrics must be reported together. Useful model FLOPs = productive time × MFU during that time × declared peak FLOPs/s, with time in seconds and the same hardware, precision and sparsity convention throughout. Use MBU for bytes against peak bandwidth, not as a substitute FLOPs factor. A running-job average that already includes stalls cannot be multiplied by ETTR without reconciling its window. → metric definitions in Chapter 0.3.

The operating metrics — what each measures and what moves it
MetricQuestion it answersFailure domainTypical range (source vintage)Primary lever to improve it
ETTRWhat fraction of scheduled-or-eligible time advanced retained progress?The job (per-run)Meta RSC large jobs: average >0.9; named populationFaster recovery: checkpoint cadence + restart time
ML goodputWhat fraction of application time advanced preserved training progress?Application; compatible fleet rollup90% vs 96% illustrative sensitivity pointsCut badput: data stalls, async checkpoint, scheduling
MFUHow much of peak FLOPs does a running job realize?The kernel/parallelism plan41%: Llama 3 405B, BF16, 16,384 H100s, 8K sequence (July 2024)Parallelism strategy, kernel/comms overlap, fabric
MBUHow much of peak memory bandwidth does decode realize?The kernel (inference)Workload-dependent; decode-boundBatching, KV-cache layout, quantization, scale-up size
AvailabilityWhat fraction of contracted time met the named service conditions?Named service boundarySite-specific; calculate from observed events and repair timesRedundancy topology, concurrent maintainability
Distinct denominators; multiply only factors proven conditional on the same ledger. Each range carries its own source vintage: the Hopper MFU example is Meta’s named Llama 3 405B configuration, and Lambda's plotted results span approximately 52.7–59.6% MFU training Llama 3.1 405B on GB300 NVL72 at 8K sequence length, depending on GPU count (November 2025). See keynumbers for sources and dates.

The table uses different denominators, and comparing two facilities on different rows is the recurring reporting error. A neocloud quoting '99.99% uptime' is on the bottom row; a customer who actually loses 20% of their training run to stragglers and restarts is living on the top row. For a contract that requires useful-output acceptance, gate production handoff on sustained NCCL/collective performance, low error rates and the agreed soak window as well as service continuity: a GPU failure costs the job its measured interruption, recovery and discarded progress. Sustained collective bandwidth, an error budget and a soak window become enforceable when the buyer names the workload, test load and pass threshold; they are not a universal 2026 SLA. → acceptance and IST in Chapter 13.6; service boundaries in Chapter 12.2.

Trace. Available allocation = 100 GPUs × 10 h = 1,000 GPU-h. Reconcile 100 + 50 + 100 + 50 + 700 = 1,000 GPU-h; discarded work is charged once. ETTR = 700/1,000 = 70%. Useful output = 700 GPU-h × 0.40 = 280 peak-equivalent GPU-h. Recovery gives 750 × 0.40 = 300; kernel tuning gives 700 × 0.42 = 294. Select recovery at the stated equal cost: 20/C versus 14/C peak-equivalent GPU-h recovered per budget unit. These are output units, not electrical energy or cash.

Flip. At equal cost, kernels win above MFU = 300/700 = 3/7 (about 42.86%); recovery must reclaim more than 700 × (0.42 − 0.40)/0.40 = 35 GPU-h to beat the stated kernel change. At unequal costs C_R and C_K, recovery wins only when C_R/C_K < 20/14 = 10/7 (about 1.43), assuming the same benefit horizon. Availability keeps its own service denominator; busy time cannot replace either account. Weight variable-width intervals by allocated GPUs. Method: Meta’s scheduled-or-eligible ETTR clock and Google’s application clock, reconciled with conditional MFU. This chapter owns the ledger; Chapter 9.4 owns recovery mathematics and Chapter 1.3/1.8 value the output.

The failure environment: why day-2 AI is not day-2 IT

Traditional enterprise IT operates a fleet of loosely-coupled, independently-failing servers: one box dies, a load balancer routes around it, and the blast radius is one request. The AI factory inverts every assumption behind that model, and the day-2 reliability program has to be rebuilt from the inverted premises.

The components fail far more often. Meta's research clusters log 6.50 failures per thousand node-days (RSC-1) and 2.34 per thousand node-days (RSC-2) — roughly 2.3–6.5 × 10⁻³ failures per node-day. In Meta's named research-cluster population, an 8-GPU job had about 47.7 days MTTF and a 1,024-GPU job about 7.9 hours; the Llama 3 405B run separately recorded 419 unplanned interruptions over 54 days on 16,384 H100s. These observations use different event populations and cannot be extended into a minute-level 131,072-GPU forecast without explicit job-membership, independence/common-mode, software, and detection assumptions. The takeaway holds regardless: failure rate rises with GPU count, so the largest jobs spend a meaningful fraction of their life recovering rather than computing. This is the arithmetic that makes recovery speed, not component MTBF, the dominant goodput lever at frontier scale. → fleet failure-rate data in Chapter 14.3; operational recovery in Chapter 14.4.

The failures are not all loud. The AI fleet has three failure classes, and the dangerous one is invisible. Hard failures (a GPU falls off the bus, a link drops) announce themselves and trigger a restart. Transient failures (a correctable ECC storm, a thermal throttle) degrade goodput without stopping the job. Silent data corruption (SDC) is the third and worst: a marginal device computes a wrong result with no error flag, quietly poisoning gradients or activations. Meta’s 22 July 2025 discussion reports about one SDC fault per thousand devices, with no exposure period from which to infer machine prevalence or a run clock. Separately, its July 2024 Llama 3 report records six SDC-attributed interruptions during a 54-day, 16,384-H100 campaign; those observations do not establish a Gemini cadence or another fleet’s rate. Soft-error susceptibility rises as devices shrink and operating voltages fall, so this is structurally getting harder — but the operating input is the named fleet's own detected and silent error rates, not a process-node comparison whose device population, exposure and error class are unstated. Detection is now a standing fleet program: Meta’s March 2022 CPU-fleet report attributes about 2.5 billion test seeds per month to Ripple, with Fleetscanner counted separately; that test volume is not GPU-fleet coverage. → SDC mechanisms and detection in Chapter 14.3.

The blast-radius problem

The defining structural feature of the AI failure environment is that the failure domain is not the failed component. In a tightly-coupled synchronous job, one GPU stalling on an all-reduce stalls the entire collective, and the whole job moves at the speed of its slowest straggler. The scale-up fabric amplifies this further: a failed NVSwitch tray degrades bandwidth for all 72 GPUs in an NVL72 domain, and an illustrative tensor-parallel group requiring all 64 GPUs, each with 99.9% same-window availability and statistically independent failures, has roughly 94% all-device availability under the model in Chapter 12.5. Correlated failures and tolerated device loss require a different model. The blast radius of a single fault is the size of the coupling domain you chose at design time, so domain sizing is a reliability decision as much as a performance one.

The cascade goes wider than the job. Meta found that 16% of total failure-related goodput loss came from secondary preemptions — small jobs getting evicted to free resources for a large job's restart. One large-job failure ripples into idle time across unrelated workloads sharing the cluster. This is the day-2 reason operators weigh the NVL72-vs-NVL36x2 fork. NVIDIA forms the two-rack configuration as one 72-GPU NVLink domain from two directly connected 36-GPU L1 domains, so a single-rack GB300 NVL72 at a 135 kW rack TDP concentrates the same domain into one rack rather than doubling its blast radius. Where scale-up domain size does change, the trade is real: a larger domain lifts the tensor-/expert-parallel ceiling and the achievable MFU and widens the blast radius of a single tray or cold-plate fault, and the right answer depends on whether you manage against MFU or against goodput. → scale-up domain sizing in Chapter 8.2; lemon-node ejection in Chapter 14.4.

Blast-radius policy — the domain-sizing fork
PolicyFailure domainGoodput on a single faultPeak MFU ceilingBest fit
One large domain (NVL72)72 GPUs / one NVLink domainWhole 72-GPU domain idles until recoveryHighest (largest TP/EP degree)Frontier dense / wide-MoE training; manage on MFU
Two 36-GPU user partitions in one NVL72 domain (admin-created)36 GPUs / user partition (logical isolation only)Workload isolation is partition-local; hardware-fault blast radius follows the failed component and topologyLower per partition (TP/EP capped at 36)Goodput-managed fleets; reliability-sensitive runs
Elastic / redundantRe-routed around the faultDegraded throughput, no full stallVariable (nonuniform parallelism)Largest runs where any full stall is unaffordable
Packaging and logical partitioning are separate: a two-rack GB200 NVL72 is one 72-GPU domain and one default partition; smaller user partitions are administrator-created.
90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

6.14 / 10.53 / 20.91%modeled
ClusterMAX April 2026 large-pretraining scenario: gold / hyperscaler / silver goodput expense
Scope & caveats

Scenario outputs for representative provider profiles, not sampled provider loss rates. Goodput expense prices extra rental time or reduced output; it includes fault-tolerance performance overhead and spare capacity as modeled. The rounded range and exact triplet describe this same scenario.

6–21%modeled
ClusterMAX April 2026 large-pretraining model: rounded goodput-expense range across provider profiles
Scope & caveats

Scenario outputs for representative provider profiles, not sampled provider loss rates. Goodput expense prices extra rental time or reduced output; it includes fault-tolerance performance overhead and spare capacity as modeled. The rounded range and exact triplet describe this same scenario.

~1.8 hr (projection)forecast
mean time to failure for a 16,384-GPU synchronous job (~7.9 hr at 1,024 GPUs)
Scope & caveats

A projection from RSC-1/RSC-2 job-failure populations, not an observation. The separately observed Llama 3 405B run on 16,384 H100s (419 interruptions in 54 days) averaged ~3 hr between interruptions with a different event population; neither extrapolates to a 131,072-GPU forecast without explicit job-membership and common-mode assumptions.

419 / 54 days
unplanned interruptions on 16,384 H100s training Llama 3 405B (~1 every 3 hr); 78% hardware-caused
Scope & caveats

Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.

The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.

~1 fault per 1,000 devices; exposure unstated
Meta reported about one SDC fault per 1,000 devices; exposure period unstated (22 July 2025)
Scope & caveats

Historical statement in Meta’s AI-hardware reliability discussion; no time denominator is supplied. Not a measured annual incidence, per-job cadence or proof that 0.1% of GPUs currently harbor a defect.

1.2–1.7%
of fleet flagged as 'lemon' nodes; ejection cut large-job failure rate 14% to 4% (+30% completion)

The reliability economics: what a goodput point is worth

Day-2 reliability is an investment with a measurable return, and the return is denominated in goodput. The canonical economic anchor: a 1 GW AI factory carries roughly $8.5B/yr in all-in TCO (Epoch AI), and at the short, frontier-economic depreciation life the figure runs higher still. Allocating that fixed annualized cost in proportion to the stipulated six-point useful-time difference gives 0.06 × $8.5B, or about $500M/yr of cost-equivalent output. This assumes the same 1 GW IT configuration, workload and cost base; it is a cost allocation, not recoverable cash. Chapter 1.3/1.8 test paid demand, marginal contribution, implementation cost and the cash-flow horizon before valuing an intervention. Underwrite each reliability intervention — checkpointing, spares, health checks, SDC scanning and node ejection — against the named fleet's measured badput reduction and implementation cost.

The economics also explain the hidden-tax structure that ClusterMAX exposed: holding GPU sticker price constant, SemiAnalysis's April 2026 ClusterMAX model puts a gold-tier neocloud's total cost below a silver-tier provider's by 5–15% on its large-training assumptions, and a hyperscaler's 36-month total swelled to 1.10x a gold-tier neocloud's — a 10% hidden tax. Modeled goodput expense is part of it (10.53% vs 6.14% for the hyperscaler and gold profiles, including their assumed fault-tolerance overhead), but SemiAnalysis attributes the gap primarily to support and setup costs (EFA performance tuning) — none of it on the rate card. The cheapest GPU-hour on paper can cost the most once delivered, because reliability is priced in goodput, not in the quoted $/GPU-hr. For fault-tolerant workloads (single-node inference), the gap collapses toward zero: the value of a goodput point is workload-dependent, and so is the reliability spend that is rational to chase it. → unit economics in Chapter 1.8.

Deep dive: reconcile ETTR before setting the recovery-time target

ETTR makes the recovery-speed economics legible when the job's scheduled-or-eligible clock is partitioned into disjoint intervals: retained productive time, eligible queueing, checkpoint stalls, discarded progress and recovery. Record detection, isolation/scheduling, reconfiguration and restore as intervals on one timeline. If two activities overlap, their union consumes wallclock once. Count discarded computation when it was first executed; replay that restores trusted progress must not be charged a second time as both a full recovery interval and an additional lost-work estimate.

Two levers still matter: shrink the checkpoint interval, which costs write bandwidth and badput unless the write overlaps compute, and shrink restart overhead, which costs hot spares and orchestration. Young's checkpoint-overhead balance and Daly's higher-order recovery treatment are derived in Chapter 9.4. Feed that method the named job's interruption distribution and measured save/restart costs from Chapter 14.4; correlated failures and checkpoint contention require the same trace to be replayed through both candidate policies. Accelerator count alone sets no two-minute target.

The operations scorecard

An operations program is only as honest as its scorecard, and the scorecard is only as honest as its denominators. The day-2 KPI set spans three layers — the workload, the fleet, and the facility — and all three belong on the scorecard with named denominators rather than collapsed into a single flattering 'uptime' figure. The workload layer is goodput and ETTR (and MFU/MBU underneath); the fleet layer is failure rate per node-day, MTTR decomposition, lemon-node rate, and SDC detection rate; the facility layer is availability, PUE/WUE, and power/thermal headroom. Each layer has a different audience: the customer cares about goodput, the reliability engineer cares about the fleet layer, and the facility team and lenders care about the bottom layer.

The recurring day-2 failure is a scorecard that reports the bottom layer (because it is the one traditional DCIM measures well) and is silent on the top layer (because it requires IT/facility telemetry correlation the legacy stack never had to do). A facility can show a perfect availability scorecard while its customers are quietly losing a fifth of their compute to badput the operator never instrumented. Closing that gap — wiring the facility telemetry to the workload telemetry so that a thermal event, a power transient, and a training stall can be correlated to a single root cause — is the central task of the observability chapter that follows. → telemetry and IT/facility correlation in Chapter 14.2.

The day-2 operations scorecard — three layers, named denominators
LayerPrimary KPIsDenominatorWho owns itWhere it is instrumented
WorkloadETTRScheduled-or-eligible GPU-timeReliability eng + customerJob eligibility and event ledger
WorkloadApplication goodputApplication elapsed GPU-timeReliability eng + customerPreserved-progress/application clock
WorkloadConditional MFU / MBUPeak FLOPs / bytes over the same productive intervalsPerformance engModel work, precision and device counters
FleetFailures/node-day, MTTR, lemon-node %, SDC rateNode-days / GPU-days in serviceSRE / fleet reliabilityHealth-checks, Fleetscanner-class scanners
FacilityAvailability, PUE, WUE, power/thermal headroomContracted time; IT energy for PUE and WUE-siteFacilities / DCIMDCIM, BMS, branch-circuit + CDU telemetry
The KPI set Part 14 instruments. The 'who owns it' column is the accountability split that prevents a metric from falling between teams.
Deep dive: lemon nodes — price ejection against lost goodput

A 'lemon' node is a machine that passes power-on and basic health checks but fails repeatedly under real workload — a marginal NVLink connector, a cold plate with intermittent flow, a GPU that throttles under sustained load. Because it boots and pings, a reachability-only availability metric counts it as 'up'; because it fails under load, it silently caps the goodput of every job unlucky enough to land on it. Meta's reliability work identified that just 1.2–1.7% of the fleet were lemon nodes, but ejecting them with >85% detection accuracy cut the large-job failure rate from 14% to 4% and improved large-job completion by ~30%.

The reason to price this intervention against other day-2 reliability spending is leverage: a tiny, identifiable fraction of the fleet causes a disproportionate share of goodput loss, and removing it requires telemetry to distinguish a load-failing node from a healthy one, orchestration to drain and eject it automatically, and capacity for testing, false-positive quarantines and qualified replacements. The dollars went not into more nines but into finding and removing the nodes that quietly destroy goodput, and the next investment must compare that useful-output gain with the facility path's maintenance, fault and contract benefit. Automated lemon-node ejection, fault isolation, and remediation are detailed as an operational discipline in Chapter 14.4; the underlying failure taxonomy and detection programs in Chapter 14.3.

Anti-patterns

The same day-2 mistakes recur, each one a consequence of reaching for the facility-availability worldview when the workload obeys goodput economics. Three are worth naming:

  • Choosing topology from checkpointability. Treating checkpoint-and-resume as proof that 2N is waste, or treating a serving label as proof that Tier-IV-class power is required. Model the named facility states, workload and fleet recovery, post-event capacity, common modes, and contract before allocating the marginal dollar. → Chapter 12.5.
  • Reporting availability and calling it goodput. A green uptime dashboard over a fleet quietly losing 15–20% of its compute to stragglers, restarts, and badput. The denominator is wallclock-of-the-facility, not accelerator-hours-on-the-critical-path, and the gap is invisible until a customer measures their own ETTR and disputes the SLA.
  • No standing SDC program. Treating silent corruption as an anomaly that conventional error alarms will catch. Meta’s fault/device statement has no exposure clock; choose screening capacity and workload checks against the measured corruption window and the cost of losing trusted progress. Without continuous scanning and checkpoint validation, multi-day rollbacks of poisoned runs are the ones the operator is structurally blind to.

Choose the next reliability intervention from useful output recovered per dollar, after the service-availability and physical-continuity gates pass. Keep occupancy, execution efficiency and retained progress as separate totals: an idle spare costs capacity, a busy GPU can repeat discarded work, and an energized hall can still miss the customer's deadline.

Serving has its own ledger. Record offered requests, accepted requests, failures, cancellations, completed outputs and elapsed observation time, with the arrival mix and admission policy. Count as useful only output meeting the declared correctness and latency SLO; include failed and timed-out requests in the service denominator rather than reporting latency only for survivors. Tokens/s needs its model, output-length and quality boundary, while GPU busy time reports resource activity. Utilization can rise as queueing violates latency. Do not multiply serving SLO attainment by an ETTR/MFU ratio unless their numerators and conditional windows actually compose. Chapter 10.6 instruments these accounts; Chapter 12.2 defines the service boundary.

This chapter sets the operating frame for all of Part 14. This chapter owns the useful-output ledger; service boundaries and continuity choices are in Chapter 12.2, with the quantitative availability/redundancy modeling in Chapter 12.5. The telemetry and IT/facility correlation that make the scorecard observable are in Chapter 14.2; the failure taxonomy, fleet failure-rate data, and SDC detection programs in Chapter 14.3; operational checkpoint/restart tuning, fault isolation, and lemon-node ejection in Chapter 14.4. The checkpoint-interval math behind ETTR is canonical in Chapter 9.4; the scale-up domain-sizing fork behind blast radius in Chapter 8.2; fleet-wide fault tolerance and autonomous recovery in Chapter 10.7; goodput-based acceptance testing in Chapter 13.6; the unit economics that price a goodput point in Chapter 1.8; and the metric definitions in Chapter 0.3.
Cite this chapter
Fehn, J. (2026). Operational KPIs, Goodput & the Reliability Economics of AI Factories (Chapter 14.1). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-1-operational-kpis-goodput-and-the-reliability-economics-of-ai-factories (accessed 2026-09-29).
@misc{aidc-14-1,
  author       = {Fehn, Jacob},
  title        = {Operational KPIs, Goodput & the Reliability Economics of AI Factories (Chapter 14.1)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-1-operational-kpis-goodput-and-the-reliability-economics-of-ai-factories},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit