Chapter 12.3
In this chapter · 6 sections
Disaster Recovery, Business Continuity & Geographic Failover
DR for an AI factory is a per-workload decision about how much spare capacity to pre-pay and where; when power-bound activation would miss the RTO, energize the spare region before the disaster and reserve surviving state, compatible GPUs and promotion authority so detection, loading and verified restoration fit the same clock.
What you'll decide here
- Which RTO/RPO tier each service actually requires, with training RPO measured back to the captured progress of the newest valid checkpoint whose copy and dependencies survive the declared disaster; underspecifying breaks the SLA and overspecifying pre-pays for continuity the workload does not require.
- Active-active across regions versus active-passive with warm/cold standby or restore-only — the fork that decides whether you keep one fleet, two, or a live fraction plus reserved activation capacity; price that readiness against the complete recovery critical path before committing the continuity budget.
- How much failover capacity you reserve and where it physically sits — if acquisition or activation exceeds the RTO, you cannot rent your way out on the day: the spare megawatts must already be energized and compatible GPUs racked. Size the surviving region for displaced peak, preemption, model residency and latency headroom.
- Which single dependencies — a region's control plane, a DNS anchor, a model-weight store, a key engineer, a single CDU vendor — can take the whole estate down at once, and which of them you are willing to leave un-hedged.
- What you have actually contracted to tenants and customers (the DR commitment in the MSA) versus what your topology can deliver under a real correlated failure — and whether your drill cadence proves the gap is closed.
Most of this Part has been about keeping a single facility running: redundancy topologies (Chapter 12.1), the goodput-versus-availability rethink (Chapter 12.2), the thermal and electrical paths that fail inside the fence. This chapter is about the failure that the fence cannot stop — the one that takes the whole site, the whole region, or the whole control plane at once. A transformer fire, a substation lost to weather, a fiber cut that isolates a campus, a cloud region's metadata anchor going dark, a wildfire-evacuation order, a ransomware event in the orchestration layer. When that happens, the question is no longer "is my UPS healthy?" It is: where does the work go, how fast, and did I pay to have somewhere for it to go?
The continuity decisions geography forces run in sequence: RTO and RPO targets set per workload class and priced by tier; active-active versus active-passive across regions, which fixes fleet count and cost; the capacity constraint that failover capacity is power-bound wherever activation cannot finish within the recovery deadline; the single dependencies (control plane, DNS, weights store, key personnel, single-vendor loops) that turn a local fault into an estate-wide one; the contractual DR obligations you owe tenants versus what the topology actually delivers; and the drills, runbooks, and FMEA tie-in that prove the plan works before you need it.
RTO, RPO, and why AI workloads split the table
Two numbers govern every continuity decision. RTO (Recovery Time Objective) is how long you can be down before recovery — the wall-clock from failure to service restored. RPO (Recovery Point Objective) is how much work you can afford to lose — the gap back to the last durable state. They are independent: you can want fast recovery (low RTO) while tolerating some lost work (looser RPO), or the reverse. The cost of each tier rises steeply and non-linearly as either target approaches zero, which is why the first act of continuity engineering is to refuse a single estate-wide target and instead tier the workloads.
AI workloads split the RTO/RPO table more sharply than any traditional enterprise estate, because the archetypes have opposite tolerances (the same split that drove the redundancy logic in Chapter 1.2 and Chapter 1.3). Interactive inference is the revenue surface and the tightest tier: a user is waiting, an SLA is running, and a regional loss that takes minutes to absorb is a visible outage. AWS's Platinum resilience example uses a 15-minute RTO and 5-minute RPO for mission-critical workloads; set project targets per workload and recovery architecture. Training continuity is checkpoint-bounded, but RPO is not the snapshot-start cadence. At failure, select the newest valid checkpoint whose durable copy and dependencies survive the event; lost work reaches back to its captured progress, not its commit-completion time; capture time, background drain, commit completion, replication scope, and overlapping checkpoints all affect that boundary. RTO is separately bounded by state discovery, capacity reacquisition, restore, validation, and resume. Batch inference can queue and retry when its delivery deadline permits; deadline-bound batch needs capacity that finishes the remaining work on time. Spending the strongest DR posture on a batch pipeline is the same anti-pattern as commissioning duplicated, fault-tolerant power when checkpoint recovery already meets the service objective and maintenance or a single component fault may interrupt load.
| Workload class | RTO target | RPO target | Failover mechanism | Cost vs single-region |
|---|---|---|---|---|
| Interactive inference | 15 min (AWS example tier) | 5 min (AWS example tier) | Active-active or hot warm standby; global load-balancer drain | Size surviving peak capacity; price the regional overhead separately |
| Training (frontier) | Selected training resume deadline | Lost-work age of the newest valid captured state that survives the event | Resume from replicated checkpoint on re-acquired GPUs | Project-measured storage, transfer and compute cost for the declared continuity design |
| Model / weights registry | Selected model-registry recovery deadline | Required surviving model/version; test replication and protected restore | Multi-region object replication; immutable versioning | Storage egress + duplicate-store cost |
| Batch inference | Contracted completion deadline or accepted deferral | Re-runnable (idempotent) | Re-queue against any region with capacity | Minimal — opportunistic spare |
| Control / orchestration plane | Selected control-plane recovery deadline | Accepted state loss under the named disaster | Multi-region quorum; no single-region anchor | Price authority, quorum, access and recovery implementation |
Zero RPO and its cost have no universal percentage. Declare the protected state and disaster first, then measure capture, blocking snapshot, background drain, durable commit, restore and contention under the reference workload. Select the newest valid checkpoint that has committed and survives the event; its captured progress determines lost work. Chapter 9.4 owns the checkpoint timing boundaries. The worked example below shows why a recent commit can still contain state older than the RPO target.
The newest valid surviving capture is S. Its state age is 09:10 − 09:00 = 10 minutes; the two minutes since remote commit are not the lost-work age. Reject the five-minute RPO claim and acquire a newer recoverable state or change the recovery objective. Flip: at the same failure time, a valid surviving capture at or after 09:05 satisfies the age screen, provided commit and recovery dependencies also survive. Retain protected older versions for corruption recovery because replication can copy deletion or corruption. PyTorch distinguishes staging and upload completion; the choice here applies the captured-state boundary from Chapter 9.4.
The spare-region problem: continuity is power-bound
What makes AI-factory DR different from a classic continuity playbook is that failover capacity cannot be conjured on demand: a tight recovery deadline turns compatible capacity into an advance procurement decision. In a traditional cloud estate, failover capacity is fungible and on-demand: a region fails, you spin up instances elsewhere, you pay the burst rate, you move on. In a power-bound AI estate, that escape hatch is closed. You cannot rent 200 MW of GPUs in a neighbouring region on the morning your primary burns, because that capacity does not exist as slack — every megawatt is contended, interconnection queues run years (the queue framing from Chapter 3.1), and a high-demand part like a current-generation accelerator is allocation-gated, not catalogue-stocked. Capacity must already be energized and racked when energization, delivery or activation would otherwise exceed the deadline. A reservation is useful only with a tested activation path.
That collapses the comforting cloud-era distinction between "reserve capacity" and "pay for it later." For the inference surface, the spare region is a real second factory, pre-paid, drawing real power, depreciating on the same contested 2–3-year bear-case economic clock as the primary (the depreciation reality of Chapter 1.3 and the economics chapter it points to). The continuity decision is therefore a capital-allocation decision, not an availability one: how many megawatts of idle-until-needed capacity will you underwrite, and can you make them earn while they wait?
The most important mitigation is to stop treating the spare as idle. Reverse-arbitrage the standby: run interruptible, RPO-loose work — batch inference, evaluation sweeps, synthetic-data generation, low-priority fine-tunes — on the failover fleet during normal operation, and pre-empt it within the tested recovery budget when the primary fails. This is the continuity analogue of the curtailable-load fast lane in power procurement: the spare region pays part of its own carry by doing displaceable work, and the failover event becomes a scheduler pre-emption rather than a cold start. The design constraint is that the pre-emptible workload must drain fast enough to hit the inference RTO — which is an orchestration property, not a hardware one.
| Posture | Standby state | Realistic RTO | Steady cost basis | Energized in advance? |
|---|---|---|---|---|
| Active-active | Live, serving, sized to absorb peer load | Detection + safe promotion + routing + SLO verification | 2× serving capacity; price total scope separately | Yes — both regions full |
| Hot / warm standby | Running, draining-ready, scaled-down | Measured promotion, eviction and scale-up | Primary fleet + scaled-down warm fraction + replication, storage, network, and licensing | Live fraction plus guaranteed activation capacity |
| Pilot light | Core/control plane up; GPU pool minimal | Measured activation and complete recovery path | Primary fleet + core recovery services + reserved scale-up capacity | Partially — depends on slack that may not exist |
| Cold standby | Defined as code; nothing running | Measured provisioning, loading and verification | Storage + IaC only | No — exposed to capacity-acquisition risk |
| Backup / restore only | Data replicated; no compute reserved | Measured reacquisition and restore path | Replication storage only | Only if reacquisition and restore meet the selected deadline |
The lower rows of the table rest on an assumption that fails exactly when it is needed. Pilot-light and cold standby look attractive because their steady cost is low — but their RTO assumes you can acquire the missing capacity when the disaster hits, and in a power-bound market that acquisition is what falls through. A cold-standby plan that depends on renting GPUs from a neocloud during a regional outage is a plan that competes for scarce capacity with every other operator whose primary just failed in the same correlated event (the wildfire, the heatwave-driven grid event, the regional storm). Cold standby is defensible for a workload that can wait through its measured acquisition and restore path — days, if that is what the path takes. An interactive endpoint with a shorter deadline needs active-active, hot standby or another prepared path that restores service in time; a long-latency batch service can buy less readiness.
For r regions, GPUs per region = ceil[18,000 / ((r−1) × 100 × 0.8)]. Two regions require 225 GPUs each, 450 total. Three require 113 each, 339 total. The one-region capacity equivalent is 225 GPUs; the ideal unrounded reserve factors are r/(r−1), or 2 and 1.5. Rounding makes the three-region factor 339/225 = about 1.5. These are capacity factors, not total-cost factors.
With three regions, normal demand is 6,000 requests/s each. After one loss, each survivor takes 9,000: 9,000/(113 × 100) = 79.6%, below the assumed 80% limit. Combined qualified capacity is 18,080 requests/s. Capacity PASS at the stated peak; a larger displaced peak or a slower destination requires a new calculation.
| Elapsed minute from disrupted service | Action / required evidence | Disposition |
|---|---|---|
| 0 | External requests fail; start engineering RTO and incident record | Clock starts |
| 1 | Independent probes detect loss; local management channel reachable | Detection complete |
| 2 | Authority declares failover; old stateful writers fenced or effect ledger enforces idempotency | Promotion authorized |
| 3 | Each survivor evicts its 38-GPU batch allocation; 75 GPUs previously served local demand | 113 GPUs available in each survivor |
| 5 | Required weights, tokenizer, adapters, keys and session recovery ready; re-prefill capacity included | Serving readiness complete |
| 7 | Routing ramp completes under surge limit; retries use logical request IDs | Displaced peak admitted |
| 9 | Agreed two-minute verification hold completes at 18,000 requests/s and reference SLO | RTO 9 minutes ≤ assumed 15-minute objective: PASS |
| Failback after repair | Reconcile state, warm primary, fence previous writer epoch and ramp back; abort on SLO regression | Separate exercise; no automatic bulk return |
Select three regions on the capacity screen: 339 installed GPUs instead of 450, saving 111 GPUs while the assumed nine-minute recovery fits the 15-minute objective. Keep the outcome on HOLD pending the qualified throughput and a drill with primary-region access denied. Flip: at 113 GPUs per surviving region, the throughput crossover is 18,000/(2×113×0.8), about 99.6 requests/s per GPU. Use the unrounded threshold for the screen; at an assumed 95 requests/s, each region needs ceil[18,000/(2×95×0.8)] = 119 GPUs, so the 113-GPU design fails. Alternatively, any additional delay above six minutes on the sequential drill path misses the RTO. AWS’s recovery-pattern guidance supplies the method boundary; Chapter 13.9 owns the recovery acceptance evidence.
Scope & caveats
Assumed 100 requests/s/GPU and 80% loading limit; equal regions, compatible GPUs, one-region-loss protection. r*ceil(18000/((r−1)*100*0.8)); capacity ratios are not cost ratios.
Scope & caveats
Named AWS example tier, not a universal practitioner target.
Scope & caveats
At failure t, select a valid checkpoint whose durable commit completed and whose storage, keys and metadata survive the declared event. Use its captured step/state time, not its commit completion timestamp, to calculate lost work.
Scope & caveats
Equal compatible regional throughput, full displaced-peak service, no second simultaneous regional loss. Storage, transfer, licenses, reservations and operating costs are priced separately.
Scope & caveats
Bulk-copy completion window, not a one-hour-RPO guarantee. Network line rate and actual inter-region transfer price require the selected path and service.
Single dependencies: how a local fault becomes an estate-wide one
Geographic failover protects against losing a place. It does nothing against losing a shared dependency — a component that, when it fails, fails everywhere at once and renders your second region useless because it depended on the same thing. Apply the blast-radius lens from Chapter 12.1 across regions: enumerate every system that is common to all sites and ask what happens when it is the thing that breaks.
A concrete warning is the AWS US-EAST-1 outage of October 19–20, 2025: a latent race in DynamoDB’s DNS-management automation produced an empty regional endpoint record. Dependent services lost connectivity and recovered on different clocks; existing EC2 instances stayed healthy while new launches failed. A multi-region data plane still depends on a single-region control plane for operations that require it. Two 2026 events add the physical-plant counterpart. Coinbase’s June 1 postmortem reports simultaneous chiller failures on May 7 in one AWS use1-az4 data hall, a thermal rack shutdown, and roughly eight hours of unavailable or degraded core services followed by further recovery: surviving zones did not automatically make the application available. Google’s July 25 report describes a July 15 3 ms utility voltage drop in europe-west4-a, a failed DRUPS transfer and a chiller controller that did not restart the pumps while redundant cooling was unavailable during construction. The hall reached 44 °C; the report summarizes 14 hours, 55 minutes of service interruption, with separate timelines for VMware Engine, Bare Metal Solution and NetApp Volumes. This is the shared-control and unavailable-reserve combination in Chapter 12.1. Keep thermal excursion, cooling restoration and service restoration on separate clocks; derive the selected liquid rack’s maintained-flow and stopped-flow response from OEM data and tests, not these air-hall timings. Your inference can span three regions, but a scheduler, discovery layer, secrets store or model-registry metadata in one place still bounds operations that need it. Test the recovery region with those primary dependencies unavailable before counting its free GPUs as usable failover capacity.
The dependencies worth enumerating for an AI estate are specific. The control / orchestration plane (scheduler, fleet-management, health-checking) — anchor it in a single region and you have re-created US-EAST-1. The model-weights store — if the failover region cannot serve the current weights because replication lagged, your spare fleet boots into a stale or empty model; weights replication must be versioned, immutable, and ahead of the failover need, not behind it. DNS and global load-balancing — the very mechanism you rely on to drain a failed region is itself a global dependency that must not share fate with it. Shared firmware and single-vendor loops — a common-cause defect in a CDU controller, a BMC firmware revision, or a single liquid-cooling vendor's pump logic can take every site running that revision down together (the thermal-path reliability concern from Chapter 12.2); this is the beta-factor / common-cause-failure problem that the quantitative model in Chapter 12.5 exists to size. And people — covered next, because the bus-factor on a novel liquid-cooled estate is realer than most plans admit.
Black-swan, pandemic, and human-continuity
Geographic failover answers "what if the site is gone." Business continuity answers the harder, slower questions: what if the people are gone, the supply chain is gone, or the operating environment changes for months rather than hours. These are the low-probability, high-consequence tails that drills rarely exercise and plans rarely fund — and that the 2020-2022 period taught a generation of operators to take seriously.
Key-personnel continuity is sharper for AI factories than for legacy halls precisely because the technology is new. The number of engineers who can safely intervene on a direct-to-chip liquid-cooled NVL72-class rack mid-fault, or who understand a specific cluster's NCCL-level failure signatures, is small — often a handful per site, sometimes one. That bus-factor is a continuity risk equal to any transformer. The mitigations are unglamorous and effective: documented runbooks that a competent on-call engineer can execute cold (not tribal knowledge), cross-training and rotation so no single person is the only path to recovery, vendor field-service contracts with guaranteed response windows for the loops you cannot self-service, and a deliberate refusal to let the on-call roster narrow to one name.
Supply continuity is the slow disaster. AI-factory recovery depends on parts that are themselves allocation-gated: HBM and advanced-packaging supply is allocated a generation ahead (Micron reported its 2026 HBM output sold out; the other suppliers publish allocation, not availability), CDUs and quick-disconnects are specialized, high-voltage transformers and switchgear carry multi-year lead times. A spare-parts strategy sized for a legacy hall — a few PSUs and fans on a shelf — does not cover a liquid-cooled estate where the failed part may be a long-lead transformer or an allocation-gated accelerator whose quoted replacement time exceeds the service restoration deadline. Continuity here means stocking the critical-spares list deliberately, holding vendor SLAs with teeth, and accepting that some failures are recovered by reconfiguration (shrinking the cluster, re-routing the fabric) rather than replacement. Pandemic / access-denial continuity generalizes the personnel question: can the site run lights-out for an extended period with no on-site staff, can remote hands and zero-touch provisioning carry the load, and is the runbook executable by people who cannot physically enter the building? The facilities that rode 2020 best were the ones already operating close to lights-out by design.
Deep dive: distributed-small-pools vs concentrated-big-site as a continuity architecture
There is a structural DR choice that sits upstream of warm-versus-cold standby, and it is geographic granularity. The instinct of a power-bound era is to concentrate — one gigawatt-class campus on one giant interconnection, because that is where the cheap firm power and the scale economics live (the siting logic of Chapter 3.1). But concentration maximizes blast radius: a single grid event, a single substation, a single weather footprint can take the entire estate. The continuity alternative is distribution: instead of one 1,000-GPU pool, consider ten pools of ~100 GPUs across mapped regional power, routing and control domains. Those counts describe partitioning, not spare capacity: size the survivors for displaced peak and test that shared dependencies cannot remove them together, using the worked case below. Failover becomes routing work-away-from-a-pool rather than standing up a cold region.
The trade is the training-versus-inference split again. Distribution is natural for inference, which is loosely coupled, latency-served, and benefits from proximity anyway — geo-distribution doubles as both DR and a latency strategy. It is hostile to frontier training, which is one tightly-coupled synchronous job that wants the largest possible single non-blocking domain and pays a convergence and bandwidth penalty for spanning regions (inter-site training is bandwidth-bound; ~1 Pbit/s inter-region targets and coast-to-coast RTT are the limiters). So the continuity architecture follows the archetype: distribute the inference surface for resilience-and-proximity, concentrate tightly coupled training where the fabric fits, then protect regional loss with a surviving checkpoint and a compatible off-site capacity path when the recovery deadline requires one. Local spares cover local failure; they do not replace campus-loss recovery. Most real estates are a hybrid — a concentrated training core plus a distributed inference mesh — and the DR plan must address each with its own posture.
Contractual continuity: what you owe vs what you can deliver
DR is an engineering posture and a set of promises in a contract, and the gap between the promise and the topology is where operators get hurt. A colocation or capacity provider's master service agreement carries continuity obligations — availability commitments, maintenance-window rules, sometimes explicit DR or geographic-redundancy clauses — and the customer's own SLA to their users sits on top. When a real correlated failure exceeds the topology, the service-credit ladder fires and, worse, the reputational and renewal damage compounds. The contract has to describe what the topology can actually deliver under a realistic failure environment, not under the optimistic independence assumption.
Three contractual failure patterns recur. First, promising an availability number the single-region topology cannot reach — committing to four or five nines on a facility whose control plane, weights store, or sole liquid-cooling vendor is a single point of failure; the availability promise breaks when that common-cause event’s frequency and complete restoration time exhaust the contracted window’s loss budget. Second, silent capacity assumptions — a DR clause that implies failover capacity exists without specifying that it is pre-energized and reserved, so that the obligation is technically met on paper and impossible in a power-bound outage. Third, mismatched RTO/RPO tiers — selling every workload the same continuity grade rather than tiering it, which either over-charges the customer for batch-grade work or under-protects their inference surface. The detailed structure of availability-versus-goodput SLAs, the penalty and service-credit ladders, and how goodput shortfalls are measured and attributed is the subject of Chapter 12.4; the productization of these commitments to customers is in Chapter 10.9 and the serving-side SLOs in Chapter 10.11. One rule connects them: never contract a continuity grade your drills have not proven.
Drills, runbooks, and the FMEA tie-in
A DR plan that has never been executed is a hypothesis. Regular, adversarial drilling converts it into a capability — the deliberate, scheduled failure of a region, a control-plane component, or a dependency, with the recovery measured against its RTO/RPO target and the gaps fixed before the real event. This is chaos-engineering discipline applied to the facility and cluster layer: you do not wait for the wildfire to discover that your weights-replication lag exceeds your inference RPO, or that the failover load-balancer shares fate with the region it is supposed to drain.
The artifacts that make drills repeatable are runbooks — step-by-step recovery procedures, written to be executed cold by a competent on-call engineer who was not in the room when the architecture was designed. A good runbook names the trigger, the decision authority (who declares a failover, since premature failover during a transient is its own outage), the exact mechanical steps, the verification that the failover succeeded, and the fail-back procedure once the primary returns — because fail-back, the step everyone forgets to rehearse, is frequently where the second outage happens. Each runbook should map to a specific failure mode in the consolidated FMEA catalog in Appendix F: every mode the FMEA enumerates as estate-significant should have a corresponding rehearsed recovery, and every drill should exercise a mode the FMEA flagged. The emergency-operating-procedures that tie operations to that same catalog live in the commissioning and operations parts; the documentation and acceptance-test discipline that captures the go-live baseline against which recovery is measured is in Chapter 13.2.
Deep dive: a minimal regional-failover runbook skeleton for an inference surface
The structure below is the irreducible skeleton of a runbook for the highest-tier case — losing a region that serves interactive inference under a named 15-minute RTO target. It is deliberately mechanical, because under real stress the value of a runbook is that it removes judgment from steps that should not require it.
1. Detect & declare. Health-checks and synthetic probes from outside the failing region trip a threshold; an on-call engineer with named authority declares the failover (a human gate prevents flapping on a transient). The RTO clock started at the service disruption, not here: timestamp detection, declaration, traffic movement, scale-up and verified restoration separately, because eight minutes of detection and decision in front of a ten-minute recovery is an eighteen-minute RTO. Any notification or claim clock the contract defines is a third set of timestamps and never a substitute for the engineering one. 2. Drain. The global load-balancer stops routing new requests to the failed region and shifts traffic to the healthy region(s) already carrying live capacity (active-active) or to the warm standby being promoted. In-flight requests are allowed to complete or are retried idempotently. 3. Promote & scale. If active-passive, pre-empt the displaceable batch work on the standby fleet, confirm the standby is serving current model weights (the version check that catches replication lag), and scale serving capacity to absorb the redirected peak. 4. Verify. Confirm latency and error-rate SLOs are met in the receiving region(s); confirm the control plane and weights store are healthy and not themselves single-region-dependent. 5. Communicate. Send the required tenant/customer notification. Measure any SLA shortfall from its defined service-impact timestamps; notification does not restart the measurement clock. 6. Fail back. Only once the primary is independently verified healthy, drain traffic back deliberately and gradually — never all at once, because a cold primary taking full load is a fresh outage. The drill that rehearses this skeleton is what turns the named RTO from a number in a contract into a number you can hit.
Choose the least costly posture whose surviving state, compatible capacity and complete drill path meet the workload’s RPO and RTO. A cheaper standby that misses either target buys a recovery promise the service cannot use; more live capacity is justified only when it removes the deadline or loss constraint. Map physical initiators to Appendix F and adversarial initiators to Chapter 11.1, then exercise the same recovery authority and evidence.
Cite this chapter
Fehn, J. (2026). Disaster Recovery, Business Continuity & Geographic Failover (Chapter 12.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-3-disaster-recovery-business-continuity-and-geographic-failover (accessed 2026-09-29).
@misc{aidc-12-3,
author = {Fehn, Jacob},
title = {Disaster Recovery, Business Continuity & Geographic Failover (Chapter 12.3)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-3-disaster-recovery-business-continuity-and-geographic-failover},
note = {Accessed 2026-09-29}
}