The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Storage & Data › 9.8

Chapter 9.8

In this chapter · 6 sections
Term help

Sizing, Data Gravity & Resilience

Size the hot tier by per-GPU bandwidth when sequential delivery binds, check metadata, retained bytes and endurance, place the corpus where the compute can reach it, and control checkpoint incast — miss a budget in the surviving system and idle accelerators absorb the cost.

POWER-BOUNDGOODPUTDENSITY-RAMP

What you'll decide here

  1. Whether the hot tier is sized to a per-GPU bandwidth target (GB/s/GPU), a capacity target (PB), or a tighter metadata, latency or endurance budget — and the capacity tier behind it — because picking the wrong primary axis strands accelerators or flash, even when the healthy system passes.
  2. How you isolate the periodic checkpoint incast from training collectives — dedicated storage rail or convergence with tested capacity and QoS — since this is a fabric-and-storage co-design problem, and sharing without a measured interference allowance leaves collective step time exposed.
  3. Whether you move data to compute or compute to data, given egress economics and gravity — the decision that quietly sets your multi-site strategy and your cloud lock-in.
  4. How much of the storage budget buys resilience (durability, multi-tenancy QoS, security) versus raw speed — and whether that resilience spend returns more as goodput than as nines.
  5. What you actually benchmark and accept on (MLPerf Storage, mixed-pipeline replay, checkpoint save/restore under load) before you sign — because datasheet peaks and pipeline reality diverge.

You now have the pieces: the parallel filesystem (Chapter 9.2), the NVMe data path (Chapter 9.3), the checkpoint math (Chapter 9.4), the loader (Chapter 9.5), the object tier (Chapter 9.6), and the inference KV hierarchy (Chapter 9.7). This chapter assembles them into a sized, sited, survivable subsystem. What makes the assembly hard is that storage failure is silent. A mis-sized fabric throws an error; a mis-sized storage tier starves the GPUs, and the loss surfaces as a utilization number a few points below where it should be, indistinguishable at a glance from a hundred other causes.

Four decisions structure what follows. Sizing: the storage:compute ratio and the bandwidth budget that decide whether the accelerators stay fed. Network co-design: how the checkpoint incast is kept off the training collectives, the placement that most affects storage goodput. Data gravity: whether you move the corpus to the compute or the compute to the corpus, and the egress economics that govern multi-site strategy. Resilience and TCO: durability, multi-tenancy QoS, security, and the benchmarking-and-acceptance regime that turns a vendor claim into a contractual commitment.

Sizing: deriving the tier from the per-GPU feed rate

The question Chapter 9.1 opened decides everything downstream: is each tier bandwidth-sized or capacity-sized? A hot tier feeding packed sequential shards or absorbing a checkpoint drain is bandwidth-sized — the governing number is GB/s per GPU, not petabytes. One serving a directory of billions of small files is metadata-sized, and one serving KV fetches is latency-sized; run the personality test from Chapter 9.1 before you name the axis. Almost every AI capacity tier is capacity-sized — it holds the corpus and the checkpoint history, and the governing number is cost per TB. Size the hot tier by capacity and you buy a vast array that cannot saturate the GPUs; size the capacity tier by bandwidth and you overpay for flash speed the cold corpus never touches. Either error strands the asset you under-prioritized.

The hot-tier budget starts from a per-GPU read target. The reference architectures give you tiers, not one floor. NVIDIA's B300 Standard tier is 80 GB/s read per 256-GPU scalable unit — 0.31 GB/s/GPU — against ~0.98 for the Enhanced tier; the DGX H200 guidelines span 0.5, 1 and 5 GB/s/GPU for Good, Better and Best, with ~4 GB/s/GPU for computer vision, with other workloads requiring their own measured read and write budgets. NVIDIA says outright that those tiers assume a mix of workloads and that you should characterize your own first. So derive the read requirement from what the loader consumes — samples or tokens per GPU-second × stored bytes per sample, adjusted for cache hit rate and concurrent jobs — multiply by GPU count, and derive the write floor from serialized checkpoint bytes and the drain window you will accept. NVIDIA's named reference tier’s published read/write pair is a cross-check on the derivation, not a replacement for it; use the exact pair rather than rounding it into a universal half-read rule. This is why the SuperPOD reference architectures publish bandwidth per scalable unit rather than per array: the B300 'Enhanced' tier lists 250 GB/s read / 124 GB/s write for NVIDIA's 32-node, 256-GPU SU, scaling to 2,000 / 992 GB/s at eight SUs, against 80 / 40 and 640 / 320 for the Standard tier. For the Enhanced definition, 250 / 256 is ~0.98 GB/s/GPU. The separate 576-GPU SU is the GB200 NVL72 rack-scale planning unit; 250 / 576 is ~0.434 GB/s/GPU, but that juxtaposes different SU definitions and is not a valid B300 normalization. You size the tier to the matching SU count, not to a capacity quote.

Capacity sizing is the other axis, governed by a small set of ratios rather than a single number. The corpus itself (raw plus tokenized plus versioned), checkpoint retention (committed generations and in-flight space times the unique serialized state from Chapter 9.4), and a working-set multiplier for shuffled access together set the petabytes. A code/text run and a frontier multimodal run can need very different corpus footprints at the same GPU count. Corpus versions, replicas and checkpoint retention drive capacity; record them before buying single- or double-digit petabytes simply because a fleet has a thousand GPUs. Size the hot tier to bandwidth and the cold tier to corpus, and don't let a capacity datasheet push you into buying the hot tier by the petabyte.

Sizing axis by tier and workload — what governs the spend
TierPrimary axisGoverning numberTypical 2026 targetFailure mode if you size on the wrong axis
Local NVMe (per-node fast tier)Bandwidth + latencyGB/s/node, checkpoint drain rateTens of GB/s/node; absorbs checkpoint burstCapacity-sized: too few drives, drain bursts saturate, checkpoints stall the step
Hot parallel FS (shared)BandwidthDelivered reads; unique checkpoint bytes / drain windowLoader-derived demand; compare with the named B300 reference tierCapacity-sized: huge array, GPUs starve at the per-GPU feed target
Capacity / object (data lake)Capacity ($/TB)PB of corpus + checkpoint retentionCorpus versions, retained checkpoints, protection and reserveBandwidth-sized: overpay for flash speed the cold corpus never uses
Inference KV tierLatency (tail) + capacityp99 read latency; GB of reusable KVFull prefix/block delivery inside the remaining SLO budgetThroughput-sized: great GB/s, but tail latency breaks the TTFT SLO
Per-GPU bandwidth targets are NVIDIA-class 2026 references; the capacity and KV rows are corpus- and workload-dependent. The point is which number governs, not a universal constant.

Follow two concurrent jobs through cold reads, overlapping saves, restore and host loss. Size the complete bid from the governing constraint: this case buys extra capacity because endurance sets the drive count.

Storage BOM case: simultaneous service and media budgets
ConstraintTrace from stated inputsResult / required action
Cold training reads768 × 20,000 × 4 / 10^9 + 256 × 500 × 0.40 / 1,0000.06144 + 51.2 ≈ 51 GB/s delivered; /0.80 ≈ 64 GB/s service floor
Concurrent drains(1,400 + 140) GB / 60 s / 0.80About 32 GB/s protected logical write service
Text restore plus running vision(1,400 GB / 60 s + 51.2 GB/s) / 0.80About 93 GB/s read service
Retained bytes300 TB + 3 × (1.4 + 0.14) TBAbout 0.30 PB logical; size from the unrounded sum
Encoded host writes/day[(1.54 TB × 86,400/270) + 16.8 TB] × 1.5 × 1.1About 840 TB/day; storage-layer maintenance is counted once
Eight-host candidate, one lost(96 − 12) × 30.72 TB × 0.30/dayAbout 770 TB/day: fails the duty
Ten-host candidate, one lost(120 − 12) × 30.72 TB × 0.30/dayAbout 1,000 TB/day: passes the assumed rating
Surviving usable capacity108 × 30.72 TB × 0.75 / 1.5About 1,700 TB logical after one loss; above the retained footprint
Displayed rounding follows scenario precision; comparisons use unrounded arithmetic. Delivered logical service, encoded host writes and wire bytes have different boundaries. Performance remains unmeasured.

Select ten hosts and 120 drives on endurance. Require 100 GB/s reads and 40 GB/s protected logical writes simultaneously with one host absent and repair active, rounding the service floors upward. Nine surviving endpoints offer 9 × 40 = 360 GB/s per direction; qualification must establish how controllers, coding and repair consume that ceiling.

The BOM has 128 compute nodes: 256 compute + 20 storage + four metadata ports = 280 endpoint links. Compute placement 43/43/42, storage 5/5/0 and metadata 1/1/0 use 49/49/42 endpoint ports plus six uplinks on each plane's leaves. All fit 64 ports. The two planes need six leaves, six spines and 36 leaf-to-spine links; each spine uses six ports. Total cables: 280 + 36 = 316; cable-end terminations: 632. All-pluggable links need 632 modules; each DAC/AOC replaces its two module positions. Chapter 8.5 owns media, lengths, spares and port modes.

Each leaf's six 200 Gb/s uplinks supply 150 GB/s line rate per direction. Five storage hosts expose 100 GB/s per plane, leaving 50 GB/s before protocol and cross-leaf repair. Placing all 32 vision nodes and ten text nodes on leaf three demands (256 × 500 × 0.40)/1,000 + (80 × 20,000 × 4)/109 ≈ 51 GB/s. Test encoded traffic against these cuts. This case covers one host loss; rack/site and plane-loss service require separate qualification.

Flip: surviving endurance permits 995.328/(1.5 × 1.1) ≈ 600 TB/day application writes, versus the assumed 509.6. Above that crossover, expand or change the qualified rating. The eight-host bid needs 840.84/(84 × 30.72) ≈ 0.33 DWPD, above 0.30. Extra namespace servers cannot repair a media-wear deficit. Complete the bid with metadata failover, firmware/power-loss behavior, GDS/client versions, licenses/support, OOB ports, racks, power, cooling and installation spares.

about 840 TB/dayderived
Encoded host-write duty in the stated storage BOM case; illustrative assumptions
Scope & caveats

Illustrative inputs; qualify project rates and limits before selection.

Network co-design: isolating checkpoint incast from the collectives

The storage subsystem does not live on its own wire. It shares — or must be deliberately kept from sharing — the fabric that carries the training collectives, and in a training cluster the most damaging storage mistakes are placement errors, not sizing errors. A checkpoint is a synchronized burst: thousands of ranks drain (write) or read at the same instant (Chapter 9.4). That is a textbook incast — many senders, converging traffic, transient congestion at the receiver's switch ports. Put that incast on the same fabric as your all-reduce, on the same cadence as your checkpoint interval, and you get a periodic collision that shows up as goodput loss correlated with checkpoint timing. The checkpoint succeeds every time; it just steals bandwidth from the collectives whenever it runs.

The placement decision has three options, and it is a co-design problem spanning storage and network rather than either alone (the fabric side is engineered in Chapter 8.5). A dedicated storage rail — a physically separate set of NICs and switches for storage traffic — means a checkpoint storm can never touch the back-end. Converging storage onto the back-end fabric with strict QoS is cheaper in NICs and switches, but you now rely on PFC/ECN priority classes and queue separation to keep the incast from starving the all-gather, and you must prove it holds under load. Sharing with no isolation works until the first concurrent checkpoint-and-collective collision, at which point step time degrades mysteriously and intermittently.

Storage-fabric placement — the isolation fork
PlacementIncremental costIsolation guaranteeOperational riskBest fit
Dedicated storage railHighest (separate NICs + switches)Physical — checkpoint incast cannot touch collectivesLowest; more cabling and ports to manageFrontier synchronous training; goodput is the headline metric
Converged + strict QoSModerate (shared NICs, priority classes)Logical — depends on PFC/ECN config holding under loadMis-tuned QoS lets incast starve all-reduce; must be load-provenCost-sensitive clusters with disciplined fabric engineering
Converged, no isolationLowestNoneHighest; intermittent goodput loss on checkpoint cadenceInference / loosely-coupled only; never frontier training
The recurring training-cluster decision. RoCE vs TCP and rail-vs-converged detail in Chapters 9.3 and 8.5. 'Cost' is incremental NIC/switch cost over a converged baseline.

Data gravity: move the data, or move the compute?

Data has gravity: the larger and more active a dataset, the harder and more expensive it is to move, and the more it pulls services and compute toward it. For an AI program this becomes a hard economic constraint at the petabyte scale, because the cloud bill for moving data — egress — is asymmetric by design. Ingress is free; egress is metered. The hyperscalers price internet egress at roughly $0.087–0.12/GB at the first paid tier (Azure ~$0.087, AWS ~$0.09, GCP Premium ~$0.12; list price, first tier — negotiated rates and the deeper volume tiers are lower), with cross-region and cross-AZ transfer adding their own per-GB charges on top. At petabyte scale these are not rounding errors: moving 1 PB out at the ~$0.09/GB first-tier rate is on the order of ~$90,000, and even at AWS's published volume tiers it is ~$54,000; a standard AI workload with weekly retraining can see egress reach 70–80% of the total cloud storage bill. Build the actual bill from stored GB-month by class, requests, retrieval bytes, each billed transfer leg, destination storage, verification and repeat reads after cache misses; AWS S3 pricing separates these components, and each candidate venue needs the same schedule on the same date.

That asymmetry forces the central question: move the data to the compute, or move the compute to the data? Move-data-to-compute is the default and is fine while the data is small or born where the GPUs are. But once the corpus is large, born in one place (a customer's region, a regulated jurisdiction, an on-prem lake), and accessed repeatedly, gravity flips the economics: bringing the accelerators to the data beats repeatedly paying egress to stream the data to a distant cluster. This is why frontier operators co-locate the prep supercomputer (Chapter 9.9) with the corpus, why regulated workloads pin compute to the jurisdiction where the data must legally remain (Chapter 10.10), and why multi-site strategy is downstream of gravity rather than the other way around.

The egress structure also creates a deliberate lock-in: because leaving costs money the incumbent keeps, gravity is a moat. The 2026 wrinkle is regulatory, and it has two separate halves that get conflated. Provider exit programs came first and are voluntary and conditional: AWS has waived egress for customers leaving the platform since March 2024, and the other majors run comparable programs you have to qualify for and apply to. The EU Data Act's outright ban on switching charges, egress included, bites only from 12 January 2027; until then providers may still levy cost-based switching fees. Either half dents the moat for a qualifying exit and does nothing for the day-to-day cross-region streaming that dominates an active AI pipeline. So design the data's birthplace and residence deliberately: land the corpus where the compute will live, replicate selectively rather than stream repeatedly, and treat every cross-region copy as a recurring egress liability, not a one-time cost.

Compare streaming and a local replica over one horizon. Streaming pays for charged cache-miss bytes on every pass plus both endpoint services. A replica pays for the initial verified move, changed-byte synchronization, protected capacity and operations. Initial move cost divided by positive recurring net savings gives the crossover horizon. Transfer bytes/sustained path payload plus verification and cutover gives the move deadline. Choose replication only if both tests pass; low reuse or high standing cost favors streaming. Data rights and recovery can exclude either option first.

Data gravity — when to move data vs move compute
ConditionMove data to computeMove compute to dataWhy
Small corpus, born near GPUsYes (default)—Egress is negligible; gravity is weak
Large corpus, accessed onceMaybe (one-time egress)MaybeSingle move may beat standing up remote compute
Large corpus, accessed repeatedly—YesRepeated egress dominates; co-locate compute with the lake
Data residency / sovereignty constraint—Yes (mandatory)Data legally cannot leave the jurisdiction (Ch. 10.10)
On-prem lake, cloud burst desiredSelective replicationHybridReplicate the hot working set; keep cold corpus put
Compare the named routes, storage classes, tiered contracts, request counts and cache misses at the same as-of date. Corpus size, repeated access and residency determine the choice.

Resilience, multi-tenancy, QoS & security

Resilience in an AI storage subsystem is a different problem from resilience in enterprise IT, because the workload values a different thing. Traditional storage optimizes durability and availability — never lose a byte, never be unreachable. An AI training storage tier optimizes goodput: keep the GPUs fed and let a failure cost as few GPU-hours as possible (Chapter 12.2). The two diverge sharply on the hot tier. A scratch/checkpoint hot tier does not need eleven-nines durability — its contents are reproducible from the last durable checkpoint and the corpus — so heavy erasure coding and cross-site replication there buy durability the workload does not value, at the cost of the write bandwidth it does. The capacity tier is the inverse: the corpus and the canonical checkpoint history are the irreplaceable assets, and that is where durability spend belongs (Chapter 9.6).

Multi-tenancy and QoS come next, and this is where a shared storage tier earns or loses its keep. A storage fabric serving many tenants or many jobs is a noisy-neighbor problem in slow motion: one tenant's metadata storm or checkpoint burst can starve another's ingestion reads, and the victim sees it as unexplained GPU idle. The controls are tenant-level bandwidth and IOPS quotas, metadata-rate fairness, and priority classes that protect the latency-sensitive (KV, ingestion) from the bursty (checkpoint). This is the storage-tier mirror of the compute-side sharing spectrum in Chapter 10.3 and the congestion engineering in Chapter 8.6. Security rounds it out: encryption at rest and in flight, tenant isolation strong enough that one tenant cannot read another's corpus or checkpoints, and — for confidential workloads — attestation-gated access consistent with the TEE model in Chapter 11.5. Weights and training data are among the most valuable and most regulated assets in the building; the storage tier is where they sit at rest.

250 / 124 GB/s
B300 Enhanced reference requirement per 32-node, 256-GPU SU; read/write, not a universal floor; Standard is 80/40 GB/s
Scope & caveats

NVIDIA B300 SU: 32 nodes, 256 GPUs; distinct from the 576-GPU GB200 NVL72 rack-scale planning SU. The Enhanced tier is 250 / 124 GB/s per SU and 2,000 / 992 GB/s at eight SUs; the Standard tier is 80 / 40 and 640 / 320. NVIDIA states these assume a mix of workloads and that requirements must be characterized per workload — they are reference tiers, not a per-GPU floor.

~14 bytes/paramderived
stated BF16 weights + FP32 master and two FP32 Adam moments; unique serialized state, before implementation metadata or replication
Scope & caveats

An arithmetic recipe for the standard BF16-compute / FP32-master / Adam mixed-precision configuration, not a measured population value. VAST Data assumed 14 bytes/param to infer model sizes from checkpoint sizes (its own footnote puts the resulting uncertainty at about ±15%); the survey did not establish the recipe. 8-bit Adam moments give 8 bytes/param; optimizers that drop the second moment change it again. Verify the tensors the chosen framework actually serializes.

$0.087/GB
Azure internet egress, first paid tier (next 10 TB after 100 GB free); list price, negotiated rates lower
Scope & caveats

List price, first paid tier; negotiated rates and deeper volume tiers are lower. Cross-region and availability-zone transfer are separate line items.

$0.09/GB
AWS data transfer out to the internet, first 10 TB/month after the 100 GB free tier; list price, negotiated rates lower
Scope & caveats

List price, first paid tier; negotiated rates and deeper volume tiers are lower. Tiers aggregate across EC2, S3 and other services; cross-region and cross-AZ transfer are separate line items.

$0.12/GiB
Google Cloud Premium Tier internet egress, first 1 TiB/month; list price, negotiated rates lower
Scope & caveats

List price, first paid tier; negotiated rates and deeper volume tiers are lower. Premium Tier only; Standard Tier and destination-specific rates (Australia, Indonesia, Korea, South America, Saudi Arabia $0.19/GiB first tier) differ.

~$90k per PBderived
to move 1 PB out at the ~$0.09/GB first-tier rate; ~$54k at AWS published volume tiers
Scope & caveats

Screening ceiling at first-tier list price; ~$54k at AWS published volume tiers; negotiated rates lower. Decimal petabyte (10^6 GB); excludes cross-region legs, requests and destination storage.

70–80%derived
source-modeled egress share for a standard weekly-retrain scenario (74% AWS / 78% Azure / 79% GCP)
Scope & caveats

cloud egress bill share

~7 days / one 512-H100 cluster
reported MTBF for one 512-H100 cluster at a top-tier operator
Scope & caveats

SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.

TCO: what storage actually costs the program

Storage is a small line on the bill of materials and a large lever on the outcome — which is exactly why it gets under-funded. The capex is real but bounded: an all-flash parallel filesystem sized to a large cluster’s per-GPU feed must be compared with the accelerator capex it keeps productive, including protected media, host service, fabric, software/support and installed power/cooling on the same duty boundary. The cost that dominates is the GPUs the storage strands. A hot tier one notch too small to keep a fleet fed, or a checkpoint path that perturbs collectives a few percent, costs goodput across the entire accelerator fleet — and on a power-bound, depreciating fleet, every point of goodput is a proportional fraction of the most expensive asset in the building. What you are pricing is the array plus the GPU-hours a wrong array burns.

The recurring costs split by tier and by venue. On-prem, the hot tier is dominated by flash media and the parallel-FS software/support; the capacity tier by $/TB and the power-and-space of dense flash or remaining HDD. In the cloud, egress and cross-region transfer can govern a repeatedly remote-read workload, as the gravity section showed, which is why a cloud storage TCO that ignores data movement is fiction. Price three things together: the media, the bandwidth (and the fabric it rides), and the movement. Optimize any one in isolation and the other two punish you.

Deep dive: benchmarking and acceptance — MLPerf Storage and the mixed-pipeline replay

The most common storage-procurement error is benchmarking the wrong primitive (Chapter 9.1). A vendor quoting 10 TB/s of sequential read has told you nothing about whether their metadata service survives a 50-million-file dataset, whether their write path sustains a checkpoint drain, or whether p99 read latency holds under a KV-offload load. Datasheet peaks and mixed-pipeline reality diverge, and that divergence is why MLPerf Storage exists. The v2.0 round (2025) drew >200 results from 26 organizations across seven countries, and — critically for this chapter — added defined checkpoint save/restore workloads against real storage, because the benchmark community recognized that the checkpoint incast is a first-class storage workload, not an afterthought. The headline finding was that tested systems served roughly twice the accelerators of the v1.0 round: storage scaling is tracking compute scaling, but only on systems that were actually designed for it. v3.0 (1 September 2026, 19 submitting organizations) pushes the suite past training and checkpointing into inference — a KV-cache test, a vector-database test, and S3 access for training, checkpointing and selected vector tests alongside POSIX — which is what you shortlist an inference or retrieval tier on now, and what finally lets an object-tier submission be read on the same axis as a file-tier one.

MLPerf is the neutral reference, but it is not your acceptance test. Acceptance is a replay of your pipeline: your file-size distribution, your shuffle pattern, your per-GPU feed target sustained across the real GPU count, and your checkpoint cadence run concurrently with synthetic collective traffic to prove the isolation holds. Make the gates contractual: (1) sustained per-GPU read at the design target across the full fleet, not a sub-cluster; (2) write drain that clears a real checkpoint within the overlap budget; (3) recovery read incast that reloads a multi-terabyte checkpoint in well under the checkpoint interval; (4) p99 latency under mixed load for any KV/inference tier; and (5) collective step-time change within the signed workload allowance when checkpoints fire. A system that passes MLPerf but fails (5) on your fabric is not commissioned. → cluster-scale benchmarking and storage/scheduler validation in Chapter 13.9.

Run the two-job workload and retain manifests, versions, topology and time series. Require cold loader rates, simultaneous 60 s drains and 5 s commits; repeat with one host absent and repair active at the 100/40 GB/s gates. In each matched repetition, loaded p99 collective step time must be ≤1.05× baseline. Above the assumed 5% crossover, choose more isolation, admission control or capacity.

Restore both jobs within Chapter 9.4's state-age, 180 s recovery and correct-next-step limits. Separately fail a namespace service and storage path against Chapter 9.2/Chapter 9.3 budgets. Test missing/corrupt generations, noisy neighbors and unauthorized reads; verify daily host writes and reserve. Procurement remains HOLD for these results and the installed power/cooling schedule. Chapter 13.9 owns commissioning execution.

Deep dive: multi-site strategy as a consequence of gravity

Multi-site is usually framed as a resilience or capacity decision; for AI it is more often a gravity decision, and getting the causality right changes the design. Three drivers force more than one site: (1) power — a single interconnection cannot energize the fleet, so capacity spills to a second campus (Chapter 3.4); (2) residency — the data legally cannot leave a jurisdiction, so compute must follow it (Chapter 10.10); and (3) proximity — inference must sit near users (Chapter 1.3). In every case the storage question is the same: where does the corpus live, and what gets replicated versus streamed?

The expensive mistake is treating multi-site as symmetric replication of everything. Replicating a petabyte-scale corpus to every site, and re-replicating every update, can make recurring egress dwarf the storage bill on charged routes. It also buys recovery reach: compare that bill with selective working-set replication and the regional recovery objective before deciding that every site needs every byte. The pattern that works is asymmetric — a canonical corpus home (where prep and the bulk of training run), selective replication of only the hot working set to satellite sites, and checkpoints written locally with a less-frequent durable copy to the canonical home for correlated-failure protection (Chapter 9.4). Cross-site failover for training is rarely worth the bandwidth; cross-site failover for inference often is, because inference is loosely coupled and latency-driven (Chapter 12.3). Decide which workload actually needs geographic redundancy before you pay to replicate the data that feeds it.

Where this is heading (2026 forward pointer)

Four trends are reshaping the sizing-and-resilience calculus, and each is treated in the consolidated roadmap (Chapter 16.2) rather than here. GPU- and DPU-initiated I/O moves the data path off the host CPU entirely: the accelerator (or the BlueField-class DPU) issues storage I/O directly, collapsing a CPU bottleneck that has capped the per-GPU feed and shifting where the storage rail terminates (Chapter 9.3). All-flash everywhere — QLC-backed object and capacity tiers displacing HDD — changes the capacity-tier TCO from a $/TB-and-spindle calculation to a $/TB-and-watt one, and narrows the bandwidth gap between hot and cold tiers (though the 2026 NAND shortage has paused the displacement on price — see Chapter 9.6). File/object convergence erodes the hard line between the parallel-FS hot tier and the object capacity tier, letting a single namespace span both and simplifying placement. And deeper CXL tiering (Chapter 9.7) inserts a memory-class tier between DRAM and NVMe that reshapes the inference KV hierarchy and, increasingly, the training scratch tier. None of these repeals the chapter's logic — you still size to a feed rate, isolate the incast, and pay for gravity. They move the numbers, not the decisions.

Choose the bid whose surviving configuration passes service, capacity, metadata and wear together, then site it using the complete movement bill and recovery boundary. The worked BOM buys more drives because endurance governs; another workload can instead buy uplinks, namespace services or a different cache policy. Accepting a paper total without the mixed-load and failure tests spends the same money while leaving the program's useful output unproven.

This chapter assembles the Part 9 components into a sized, sited, survivable subsystem. The per-GPU bandwidth framing and the bandwidth-vs-capacity fork originate in Chapter 9.1; the parallel filesystem that serves the hot tier is Chapter 9.2; the CPU-bypass data path and storage-rail placement is Chapter 9.3; the checkpoint math that sizes the incast is canonical in Chapter 9.4; the loader path in Chapter 9.5; the object/capacity tier in Chapter 9.6; the inference KV hierarchy in Chapter 9.7; and the prep supercomputer that gravity co-locates with the corpus in Chapter 9.9. The fabric isolation that keeps checkpoint incast off the collectives is engineered in Chapter 8.5 and Chapter 8.6; multi-tenancy and QoS mirror the compute-side spectrum in Chapter 10.3; storage security and confidential access tie to Chapter 11.5; the goodput-vs-availability reframing that makes resilience a GPU-efficiency problem is Chapter 12.2; geographic failover is Chapter 12.3; storage acceptance and cluster-scale benchmarking is Chapter 13.9; the energy-supply driver of multi-site is Chapter 3.4; the legal-residency driver is Chapter 10.10; and the 2026 storage roadmap is consolidated in Chapter 16.2.
Cite this chapter
Fehn, J. (2026). Sizing, Data Gravity & Resilience (Chapter 9.8). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-9-storage-and-data/9-8-sizing-data-gravity-and-resilience (accessed 2026-09-29).
@misc{aidc-9-8,
  author       = {Fehn, Jacob},
  title        = {Sizing, Data Gravity & Resilience (Chapter 9.8)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-9-storage-and-data/9-8-sizing-data-gravity-and-resilience},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit