Guide › Part 12
Part 12
Reliability, Resilience & Standards
5 chapters
12.112.212.312.412.5
Resilience Standards, Redundancy Topologies & Fault-Domain Engineering
Tier and Rated standards certify facility topology only; an AI operator must separately choose fault domains, blast-radius limits, and whether to buy concurrent maintainability, fault tolerance, or both.
The AI-Cluster Reliability Rethink: Goodput vs Facility Availability
An AI cluster earns its return on goodput — the fraction of bought GPU-hours doing useful work — and for most AI factories the next redundancy dollar buys more goodput than facility nines.
Disaster Recovery, Business Continuity & Geographic Failover
DR for an AI factory is a per-workload decision about how much spare capacity to pre-pay and where; because GPU capacity is power-bound, the spare region must be energized before the disaster.
SLAs, Goodput Contracts & Availability Commitments
Promise facility availability to a tenant who is paying for goodput and the SLA tracks neither the customer's pain nor the provider's control; match the contracted metric to the workload.
Quantitative Reliability & Availability Modeling (RBD / FTA / Monte-Carlo)
RBD, fault trees, Markov chains, and Monte-Carlo turn the redundancy debate into a model in which every nine — and every percent of goodput — has a traceable parent in component failure rates.