Chapter 12.5
In this chapter · 9 sections
Quantitative Reliability & Availability Modeling (RBD / FTA / Monte-Carlo)
RBD, fault trees, Markov chains and Monte-Carlo turn the redundancy debate into a model in which every availability result and percent of goodput has traceable event, exposure and recovery inputs; use one event inventory, reproduce the loss distributions and price the uncertainty before choosing the investment.
What you'll decide here
- Which target you are actually modeling — facility availability (uptime fraction) or cluster goodput (effective-training-time fraction) — because they are different objective functions with different dominant terms, and a redundancy investment that buys nines may buy almost no goodput.
- Which method fits the question: closed-form RBD/k-of-n algebra for static redundancy, a Markov state-space model when repair-crew limits and degraded states matter, fault trees and minimal cut sets to find the dominant failure path, and Monte-Carlo when the failure environment is correlated, time-varying, or non-exponential.
- Whether the model catches common-cause failure — the shared CDU, shared bus or fleet-wide firmware push that can erase the apparent protection of parallel units — because an idealized parallel block claims six nines that a 5% beta-factor erases down to roughly four; represent known events explicitly and reserve beta for residual causes so the same loss is never counted twice.
- Which inputs you trust: the Chapter 14.3 AFRs, the Appendix F FMEA scenarios, and the repair/restore times — because a quantitative model is exactly as credible as its rates, and most of those rates are contested and scale-dependent.
- Where the next dollar of redundancy buys the most nines or the most goodput — the sensitivity (importance) ranking that is the actual deliverable of this chapter and the engine under the Chapter 12.2 tradeoff curve.
Chapter 0.5 gave you the vocabulary — series and parallel, N / N+1 / 2N, the Tier ladder, availability versus goodput — as a primer you could carry into a design review. Here that vocabulary becomes a method: quantitative machinery that takes component failure rates and repair times and rolls them up into a defensible cluster-level number, with a traceable event, exposure and recovery parent for every numerical availability target and every percent of goodput. The worked case supplies an executable script, a complete synthetic ledger and a sensitivity calculation. "We built it 2N, so it's reliable" stops being an assertion and becomes a calculation someone can check, attack, and improve.
The techniques stack. The series/parallel/k-of-n algebra extends into real availability arithmetic; Markov state-space models handle repairable systems where repair crews are finite and degraded operation is a real state, not a binary; fault trees and minimal cut sets surface the dominant failure path that the redundancy diagram hides; beta-factors put a number on common-cause failure, which is where idealized parallel redundancy quietly fails to deliver; and Monte-Carlo takes over once the failure environment becomes correlated, time-varying, and non-exponential — when the service’s observations require those distributions. A model is exactly as good as its inputs and its treatment of correlation, and the two most common ways an availability model misleads are assuming independence and modeling uptime when the thing that pays the bills is goodput.
Two targets, two objective functions
Before any algebra, name the target. A traditional IT availability model computes the fraction of time the service is up — the complement of downtime under the selected state definition. A contractual uptime floor is numerical; Tier requirements concern assessed infrastructure outcomes. An AI training cluster cares about a different number: goodput, the fraction of wall-clock GPU-time that advances the job rather than being lost to failure, detection, restart, and recompute-from-checkpoint. The two are not the same function, and the gap between them is the entire reason Chapter 12.2 exists. A facility can be 99.99% available while its cluster achieves 85% goodput — illustrative percentages for the same resource and time boundary, with GPU failures, checkpoint restarts and replay eating progress while power and cooling mostly stay present. The models have different loss terms, but shared facility events still enter both calculations.
Each investment moves a different set of states. Duplicated facility power can prevent or ride through modeled source and distribution-path events; it does not prevent GPU, HBM, optic, software, or checkpoint failures. Checkpointing, hot spares, and lemon-node ejection reduce workload losses; they do not keep a hall energized or prove that its recovery stays inside the contract. A defensible model declares the target and state set, then compares incremental lifecycle cost with the frequency and consequence each mechanism removes. → Chapter 12.2.
Reliability block diagrams: the series/parallel/k-of-n algebra
The RBD is the workhorse and the right place to start, because it makes the topology's logic explicit before any number is plugged in. Each block has long-run availability A = mean uptime / (mean uptime + mean restoration time). State what the source calls MTBF: a failure-to-failure interval that already includes repair must not have repair added again. Independent block states compose by two rules; these product identities do not require exponential lifetimes. Series (every block must work): availabilities multiply, A = A₁ × A₂ × …, so a chain is always less available than its weakest link and unavailability roughly adds. Parallel (any one block suffices): the unavailabilities multiply, A = 1 − (1−A₁)(1−A₂)…, so redundancy multiplies the small failure probabilities into a much smaller one — this is where the nines come from.
The general case is k-of-n: the system works if at least k of n identical units work, which is the honest model for an N+1 or N+2 power or cooling plant (you need N running, you installed N+1, so it survives any one failure). For independent units each with availability p, the system availability is A = Σ(j=k…n) (n choose j) pʲ(1−p)ⁿ⁻ʲ, the probability that k or more of n survive. Shared crews require the joint-state probabilities instead. The practical consequence: in the high-availability independent-unit regime, moving from N (k=n, all must work) to N+1 (k=n−1 for that installed set) buys a large drop in capacity-loss probability, while further spares yield diminishing absolute reductions; compare those purchases with faster repair and removal of shared causes. For independent identical units at high availability, that diminishing-returns curve is not intuition; it falls straight out of the binomial, and it supplies the loss-probability side of the Chapter 12.2 tradeoff curve.
| Method | Answers | Key assumption | Breaks when | Typical AI-DC use |
|---|---|---|---|---|
| RBD / k-of-n algebra | Steady-state availability of a fixed topology | Independent block states; each block has the appropriate availability for the chosen time basis | Repair is shared/limited, or failures correlate | Power & cooling plant nines; N vs N+1 vs 2N |
| Markov / state-space | Repairable systems with degraded states & finite crews | Exponential (memoryless) transitions | Wear-out, time-varying, or non-exponential repair | MTTR vs crew size; degraded-cooling ride-through |
| Fault tree / minimal cut sets | The dominant failure path and its top contributors | Boolean gates; rates assignable to basic events | Dynamic ordering / sequence-dependent failures | Finding the single shared CDU or bus that dominates |
| Monte-Carlo simulation | Availability AND goodput under a realistic environment | Only what you put in the sampling model | Inputs (rates, correlations) are themselves wrong | Cluster goodput with correlated, scale-dependent AFRs |
The table is an escalation ladder rather than a menu. Start with the RBD because it is cheap and forces the topology to be explicit. Escalate to Markov the moment repair is shared or the system has a meaningful degraded state — a cooling plant that limps at reduced capacity rather than failing outright. Build a fault tree when the RBD has gotten complex enough that the dominant failure path is no longer obvious by inspection. And reach for Monte-Carlo when the thing you actually care about — the distribution of goodput under a correlated, scale-dependent failure environment — outruns the closed forms; an analytical expected-ETTR still makes a useful check on the simulation's mean. Each rung exists because the rung below it made an assumption that the cluster violates.
Markov state-space models: repair crews and degraded states
An independent-unit RBD often hides a repair assumption: every failed unit gets its own crew when it fails. A shared crew couples those repair states and breaks the product, though an RBD can still represent the whole pool as one block using its joint-state availability. Real facilities have finite crews, parts depots with turnaround time, and components that fail into a degraded state rather than a clean binary down. A Markov model captures all three by representing the system as a set of discrete states (all-up, one-down-repairing, two-down, degraded-running) with transition rates between them: failure rates (λ) push the system toward worse states, repair rates (µ) pull it back. For n equal units and one crew, let i count failed units: i→i+1 has rate (n−i)λ and i→i−1 has rate μ for i>0. With independent unlimited crews the latter rate is iμ. Solve stationary weights w₀=1, wᵢ=wᵢ₋₁(n−i+1)λ/μ, normalize to πᵢ, and sum acceptable states n−i≥k — and, critically, you can see how much of your unavailability is queueing for a repair crew rather than waiting on a part.
This is where MTTR stops being a single number and becomes a policy. A 2N plant whose second unit is down for 24 hours awaiting a part is, for that window, running N with no margin — an operating state hidden when the RBD is reduced to a single availability scalar. The Markov model prices that exposure: if the repair rate µ is slow relative to the failure rate λ, the probability of sitting in the vulnerable degraded state climbs, and the second-failure-during-repair term — the one that actually takes the plant down — grows with it. The decision this surfaces is concrete and expensive: does the next dollar buy another redundant unit, or a faster repair (on-site spares, a larger crew, a depot SLA)? For a plant already at N+1, shrinking MTTR often buys more availability per dollar than adding N+2, because it attacks the dominant second-failure-during-repair term directly. The model tells you which; intuition routinely gets it backwards. → repair-time and sparing inputs in Chapter 14.3.
Deep dive: when memoryless transitions become a convenient lie about age
Markov models assume memoryless (exponential) transitions: the probability a component fails in the next hour does not depend on how long it has already run. That is a genuine convenience — it makes the math a linear system you can solve by hand or with a small solver — and it is roughly true during the flat-bottom of the bathtub curve, the useful-life region where failures are random. It is false at both ends. During infant mortality (the first weeks after install, the burn-in window where new GPU clusters fail far more often than mature ones), and during wear-out (aging fans, pumps, capacitors, optics with creeping link margin), the failure rate is not constant and a mature-fleet exponential fit understates risk when the new cohort’s hazard is higher; inspect age and exposure, because another fitted window can err in the other direction.
For an AI cluster this is not academic, because a capacity ramp keeps part of the fleet new — a freshly installed tranche has not yet demonstrated the mature fleet’s failure rate. The fix when the rates differ is one of two things: piecewise-Markov (different rate regimes for burn-in, useful-life, wear-out) when age bands fit the observations, or Monte-Carlo sampling from Weibull or empirical lifetime distributions when the observed age dependence needs that detail. The tell that you have outgrown the exponential assumption is a model that is confidently wrong about a brand-new cluster’s first operating window; validate that cohort separately before using a mature-fleet fit to promise its goodput.
Fault trees and minimal cut sets: finding the dominant path
An RBD is built bottom-up from components; a fault tree is built top-down from the thing you are trying to prevent — "cluster job halted," "hall loses cooling," "tenant SLA breached" — decomposed through AND/OR gates down to basic events with assignable rates. The two are duals, but the fault tree answers a question the RBD does not phrase well: what are the smallest sets of simultaneous failures that take the system down? Those are the minimal cut sets, and ranking them by probability is the output practitioners reach for first. A first-order cut set (a single basic event that alone causes the top event) is a latent single point of failure hiding inside a topology you thought was redundant. A model that surfaces one shared, un-redundant component buried under two layers of N+1 has earned its keep. An RBD can represent that element too; the fault tree’s job is to challenge whether the event inventory left it out.
The recurring AI-data-center example is the cooling distribution unit. The RBD shows N+1 CDUs and reports comfortable nines. The fault tree, traced to basic events, reveals that all of them share a single secondary-loop isolation valve, a single facility-water supply, or a single controls PLC — a first-order cut set that the parallel block diagram drew right over. The same pattern recurs for a shared busway feeding ostensibly-independent power paths, a single BMS that gates every chiller, and a single firmware image across every BMC. Push the tree down to basic events — explicit hardware, software, control or procedural events with a defined occurrence and recovery basis — rather than stopping at the convenient block boundary, because the convenient boundary is exactly where the shared dependency hides. → the FMEA basic-event catalog in Appendix F.
Common-cause failure: the beta-factor that erases your nines
This is the term that most often separates an overstated availability number from a defensible one. Parallel redundancy delivers its spectacular nines only if the redundant units fail independently. They never fully do. A shared cause — a common firmware bug pushed fleet-wide, a shared coolant chemistry problem, a single controls fault, a power transient that hits every unit at once, a maintenance error repeated on each "redundant" path — defeats redundancy by failing the units together. The standard way to model this is the beta-factor: a fraction β of a component's failures are assumed common-cause, hitting all redundant units simultaneously, while the remaining (1−β) fail independently. IEC 61508-6:2010 Annex D maps hardware common-cause scoring to β bands of 0.5% to 5% for logic subsystems and 1% to 10% for sensors and final elements, as described in SINTEF’s 2015 common-cause study. Use the applicable licensed scoring procedure for the named subsystem; these bands are not a generic data-center beta or measured fleet failure rate.
The consequence is severe and counterintuitive. An idealized parallel pair of units each at 99.9% availability computes to about six nines (1 − 0.001²). Layer in a 5% beta-factor — assume 5% of failures are common-cause — and the system unavailability becomes β·U + (1−β)²·U² ≈ 0.05 × 0.001 + 0.9 × 10⁻⁶ ≈ 5.1 × 10⁻⁵: roughly four-and-a-quarter nines, because the common-cause term, not the independent-failure product, now dominates. Where the shared event is named — a controls PLC with its own failure rate λs and restoration rate μs — model it explicitly as a series element with availability μs/(λs+μs), specify its recovery distribution, and reserve β for the causes still unrepresented; β alone cannot convert component unavailability into a service result. Redundancy cannot buy you below the common-cause floor. This is why diversity is worth paying for: different firmware versions across redundant controllers, A/B coolant sourcing, independent power-transient paths, staggered maintenance windows, and, above all, never pushing the same firmware to every redundant unit on the same day. A model that omits shared causes can claim redundancy nines those causes erase; one that represents the shared events explicitly does not need to add beta again for the same failures.
Monte-Carlo: simulating availability AND goodput under a real failure environment
Closed-form methods buy tractability with assumptions the cluster violates: independence, constant rates, exponential repair, a single target. Monte-Carlo trades a compact algebraic answer for repeated samples of the declared event and recovery process; its honesty still depends on what those samples represent. You build a generative model — sample each component's time-to-failure from its real lifetime distribution (Weibull for wear-out, empirical for the burn-in hump), inject correlation explicitly (a common-cause event that fails a set of units together, a grid transient that hits the whole hall), model the repair queue with its finite crews and depot delays — then run the clock thousands of times and read off the distribution of outcomes, not just a point estimate. The output is a distribution of annual downtime, goodput and bad windows, with confidence intervals for finite-sample error kept separate from uncertainty in the input rates. Analytical methods can also produce distributions where their assumptions fit.
For an AI cluster the killer feature is that the same simulation produces both targets at once. Layer the workload model on top of the failure model — checkpoint interval, detection time, restart-and-recompute cost, lemon-node ejection policy — and the run yields goodput directly: every sampled GPU failure triggers a detection delay, a restart, and a recompute of the work since the last checkpoint, and the fraction of GPU-time that survives that gauntlet is the goodput. This is what you reach for when the question is "if I halve my checkpoint interval, what happens to my goodput distribution at 100k GPUs?" — because the answer depends on a failure rate that scales with cluster size, a recompute cost that depends on interval, and a correlation structure the simple algebra drops. Simulation is not the only honest route: Meta's closed-form expected-ETTR result gives a cheap baseline for the mean, and renewal or state-space treatments carry much of the rest. What simulation buys is the distribution and its tail, which an expectation cannot give you. The cost is calibration: a Monte-Carlo result is only as good as the rates and correlations you fed it, which is why this method consumes the Chapter 14.3 AFRs and the Appendix F scenarios as its raw inputs rather than inventing them.
Scope & caveats
Two A100 research clusters with distinct workloads; failures divided by sum of job runtime × allocated nodes. Not unique failed-FRU counts or a universally portable component failure intensity.
Scope & caveats
Projection from a failure model fitted to 11 months of two A100 research clusters (RSC-1/RSC-2, >150M A100 GPU-hours). The paper projects 1.8 hours at 16,384 GPUs and 0.23 hours at 131,072 GPUs. Not a measurement, not an H100/Blackwell-era observation, and distinct from Meta's separately published 16,384-H100 Llama 3 run.
Scope & caveats
Hypothetical dedicated RSC-1 case with nonblocking checkpoint writes: 60-minute versus 5-minute interval. Not a production observation or a universal improvement curve.
Scope & caveats
Named RSC deployment and event definition; a useful external comparison, not a required calibration target for other fleets.
Scope & caveats
Hardware common-cause bands in the IEC functional-safety context; no portable data-center β. Score the subsystem with the Annex D checklist before using a band.
Scope & caveats
Job-interruption event counts for one named run. The source does not establish unique failed devices or equipment population-time exposure, so these counts must not be annualized into component AFR, fleet lambda, cumulative equipment risk or spares demand.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
The roll-up: from component AFRs to cluster availability and goodput
For the CDU pool, λ/μ = 0.004 gives stationary weights [1, 0.012, 0.000096, 0.000000384]; the acceptable numerator is 1+0.012. For the power pool, weights are [1, 0.008, 0.000032] and the acceptable numerator is 1+0.008. Normalize each by its full sum. Shared-control availability is 0.5/(0.5+0.0005). Multiplying the independent pool availabilities gives expected annual downtime (1−A)×8,760, about 9.9 hours/year. The top event is S OR (P1 AND P2) OR any CDU pair; minimal cuts are {S}, {P1,P2}, {C1,C2}, {C1,C3}, {C2,C3}. Removing S leaves about 1.1 hours/year from the two equipment pools.
The script’s reduced node-only renewal check gives about 90.9% goodput; it imports the compute/save/recovery timing contract from Chapter 9.4. NIST’s r-out-of-n method supplies the independent-state comparison; the shared-crew recurrence above is the guide’s explicit state extension.
Runnable analytical checks for the illustrative model
Run this Python standard-library script to reproduce the pool-state availability and reduced node-only goodput check.
from math import expm1
def pool(n, k, lam, mu):
w = [1.0]
for i in range(1, n+1):
w.append(w[-1]*(n-i+1)*lam/mu)
return sum(w[:n-k+1])/sum(w)
power = pool(2, 1, 1/2000, 1/8)
cooling = pool(3, 2, 1/1000, 1/4)
shared = (1/2)/(1/2+1/2000)
a = power*cooling*shared
print('Facility downtime h/year:', round((1-a)*8760, 1))
print('Without shared event h/year:', round((1-power*cooling)*8760, 1))
# Reduced node-only check using the Chapter 9.4 timing boundary.
t, save, restore, lam = .5, 1/60, .5, 1/12
node_goodput = t/(expm1(lam*(t+save))*(1/lam+restore))
print('Node-only goodput %:', round(100*node_goodput, 1))
assert abs(a-0.9988741518) < 1e-10
assert pool(3, 2, 0, 1/4) == 1
The point of the method is the roll-up — taking the per-component failure rates from Chapter 14.3 and the failure scenarios from Appendix F and composing them, level by level, into a number for the whole cluster. The path runs: component AFRs → per-node failure rate → RBD/k-of-n for the redundant subsystems (power, cooling, network core) → fault tree to catch the shared cut sets → renewal, state-reward or event simulation for synchronous-job goodput → cluster-level availability AND goodput. Each level has a method matched to its question, and the levels compose: the facility’s event rates, durations and affected load feed the goodput process — a power event that interrupts required ranks is another correlated failure that stalls the job — while mean availability alone cannot identify the recovery cost, and population-time equipment failure estimates derived from unique FRU events become one input to the goodput run; Meta's job-interruption counts do not supply component AFR or per-node λ.
Check the roll-up against ground truth before trusting it forward. Validate a goodput model against the named fleet/job's observation window and badput decomposition. Use 90% and 96% only as stipulated guide sensitivity endpoints; calibrate against the named fleet/job’s measured event ledger; show sensitivity to event definition, workload mix, checkpoint policy, and correlated failures, and do not tune failure rates merely to reproduce that pair. Verify that the modeled state behavior satisfies the selected Tier outcome, then calibrate numerical availability against the named facility's incident and maintenance history or a declared empirical dataset. The model earns field use when its number survives both comparison with relevant published benchmarks and prediction of an untouched incident-history window with matching event definitions, exposure and workload boundaries. The worked case below proves arithmetic, not field calibration. → AFR inputs and the failure taxonomy in Chapter 14.3; the FMEA scenario catalog in Appendix F; the checkpoint math behind the goodput term in Chapter 9.4.
Sensitivity analysis: where the next dollar buys the most nines or goodput
The deliverable is a ranking, not a single availability number. Once the model is built and calibrated, you compute the sensitivity (importance) of the result to each input: how much does cluster availability or goodput move per unit of investment in this component, this redundancy step, this MTTR reduction, this checkpoint interval? That ranking is the engine under the Chapter 12.2 tradeoff curve and the answer to the only question the CFO actually asked: where does the next dollar of redundancy buy the most?
The rankings routinely surprise. On the availability side, shrinking MTTR with on-site spares, a depot SLA or a larger crew attacks the second-failure-during-repair window directly and can beat another redundant unit. Removing a shared cause can beat both when that event dominates. Compare all three in the same state process and cost horizon; the worked prices below select positive incremental value, and changed shared-event exposure can reverse the choice. On the goodput side, Meta's hypothetical RSC-1 scenario shows how far checkpoint cadence and node policy move goodput: moving a 16k-GPU run from a 60-minute to a 5-minute checkpoint interval lifts expected ETTR from 0.70 to 0.93, while its deployed lemon-node ejection cuts 512+-GPU job failures from 14% to 4%. Use the model's sensitivity ranking to spend on the selected service's dominant interruption and recovery states — checkpoint and node policy in this scenario, but liquid-flow continuity, replica placement, or facility path states in others. The sensitivity analysis is what proves that to a skeptic with a budget, and it is why this model, not intuition, should set the redundancy spend. → the tradeoff curve it feeds in Chapter 12.2; the SLA commitments it underwrites in Chapter 12.4.
| Change | Mean annual goodput | Recovered GPU-hours/year | Net value/year at $2.0/GPU-hour |
|---|---|---|---|
| Base | 90.8% | — | — |
| Fast checkpoint/restore | 96.3% | about 500,000 | about +$800,000 |
| Control improvement | 90.9% | about 9,800 | about +$10,000 |
| Fourth CDU | 90.8% | about 1,200 | about −$48,000 |
| Faster equipment repair | 90.8% | about 1,100 | about −$23,000 |
| Third power path | 90.8% | about 720 | about −$150,000 |
Select fast recovery under these assumptions: it recovers about 500,000 GPU-hours/year and has the largest positive net value. The combined fast-plus-control run reaches about 96.4% goodput and adds about $9,000/year net over fast alone; select both at these prices. Recompute that marginal value after any input change. The additional equipment options fail this discretionary economic screen at their assumed prices; they remain mandatory if needed for a required state. Flip: fast recovery ceases to earn its cost when its annual increment exceeds ΔG×1,024×8,760×$2.0, about $1.0 million/year at the base rates. Increasing shared-control event frequency tenfold lowers base goodput to about 89.7%; buying more independent units cannot remove that shared-event loss.
The base annual mean is about 90.8%; the simulated annual fifth-to-95th-percentile range is about 90.4%–91.2%. The simulation reports its own sampling error on the mean; the runnable script above reproduces only the analytical checks. Neither that sampling error nor the window percentiles measure uncertainty in the assumed rates. Under Exhibit T’s hypothetical 92% floor, the separate base 720-hour simulation breaches in 1,880 of 2,000 windows with about 12.8% mean credit exposure; fast recovery has zero breaches in this sample. Zero simulated breaches is not zero risk. Vary rates and recovery inputs before pricing a monthly promise, then validate on a held-out fleet window with matching exposure definitions.
Scope & caveats
All ledger and mitigation rates, timing, 1024-GPU population, annual costs and $2.0/GPU-hour value are assumed. Prices: 200000 fast, 10000 control, 50000 fourth CDU, 25000 repair, 150000 third path, 210000 combined USD/year. 2000 windows, seeds 120500+i.
Deep dive: the reliability model and design-validation twin
The reliability model uses event rates, recovery durations and state probabilities to price interruptions and retained work over time. The physics-based twin uses geometry, materials and boundary conditions to test temperatures, pressures and transients. A twin’s qualified degraded-flow limit defines an acceptable reliability state; a failure model estimates time spent there. Neither result substitutes for the other. Chapter 2.7 owns the design-validation twin and its physical derivations.
Anti-patterns
The same modeling errors recur, each one a way of producing a number that is precise and wrong. Four are worth naming:
- Using the workload metric as a topology selector. A training goodput model that omits facility maintenance and fault states can under-provide continuity; a facility RBD that omits silicon, checkpoint, and recovery losses can overstate delivered value. Model both, then compare N, N+1, distributed, or 2N with compute-stack and fleet alternatives. → Chapter 12.2.
- Assuming independence. Reporting the parallel-product nines without a beta-factor — six nines that a 5% beta collapses to roughly four — while omitting the shared bus, CDU loop, controller or fleet-wide firmware push that can fail every ‘redundant’ path together. An essential shared series element caps the block at its own availability; model its rate and recovery, plus repair coupling, and use beta only for residual causes not already counted. The most common way an availability number is overstated.
- Stopping the fault tree at the convenient block. Drawing N+1 CDUs as independent parallel blocks without pushing to basic events, and missing the single shared valve or PLC that is a first-order cut set. The redundancy diagram drew right over the single point of failure. → Appendix F.
- Modeling once and reusing across scale. Taking the goodput number from a 16k-GPU model and assuming it holds at 100k, when independent equal per-node faults make a restart-all job’s MTTF fall roughly inversely with node count, while changed recovery algorithms, placement and shared events can alter that scaling. The model must be re-run at each point on the ramp. → Chapter 14.3.
Acquire the event counts, operating exposure and complete recovery traces from Chapter 14.3 before committing to the modeled service gain; guessed rates cannot support that promise. Appendix F owns physical failure-mode identifiers; Chapter 11.1 owns adversarial initiators. This chapter owns their stochastic consequences.
Cite this chapter
Fehn, J. (2026). Quantitative Reliability & Availability Modeling (RBD / FTA / Monte-Carlo) (Chapter 12.5). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-5-quantitative-reliability-and-availability-modeling-rbd-fta-monte-carlo (accessed 2026-09-29).
@misc{aidc-12-5,
author = {Fehn, Jacob},
title = {Quantitative Reliability & Availability Modeling (RBD / FTA / Monte-Carlo) (Chapter 12.5)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-5-quantitative-reliability-and-availability-modeling-rbd-fta-monte-carlo},
note = {Accessed 2026-09-29}
}