Chapter 1.2
In this chapter · 8 sections
Training Data Centers: Synchronous, Dense, Checkpointable
A training cluster is one synchronous supercomputer moving at its slowest GPU's pace; design for goodput per megawatt—dense and checkpointable, with cooling and fabric selected from the named rack and measured traffic/SLO requirements—before you cut steel.
What you'll decide here
- The scale-up domain size (the cited 8 / 72 / 576-GPU profiles have different product and roadmap boundaries) you build to — because it sets your tensor- and expert-parallel ceiling, your in-rack copper budget, and how much traffic you keep off the expensive scale-out fabric.
- What blocking ratio each back-end tier can support under the measured collective traffic, placement, communication overlap, failure headroom, and step-time/MFU target—and how target-fabric validation proves it.
- The density tier you plumb for now (132 kW NVL72 → 600 kW Kyber) versus the IT you fit out today — price the irreversible substrate (floor, water, electrical headroom) against each named ramp scenario and release only the funded reserve.
- The reliability posture: compare N, N+1 and 2N with disciplined checkpointing against the same maintenance, fault and recovery requirements; checkpointability prices lost work but does not prove that a path may interrupt.
- Whether this is one campus or a multi-site / gigawatt run — which forces a choice between synchronous cross-DC fabric and asynchronous (DiLoCo-class) training, and rewrites the inter-campus fiber and power-resilience plan.
Meta's Llama 3 405B run held 16,384 H100s in lock-step on a single job for 54 days, the whole run moving at the speed of its slowest straggler. That is what a pre-training cluster is: one machine — a single tightly-coupled supercomputer — not a fleet of independent servers that happen to share a building. That one fact reorganizes every downstream decision. The design objective is goodput — the fraction of wall-clock time the cluster spends doing useful gradient work — measured against the megawatts you can energize and the depreciation clock already running on the silicon; the facility must also satisfy the declared maintenance and fault requirements. Everything in this chapter follows from optimizing goodput per megawatt on a job that cannot tolerate a slow link, a hot GPU, or an unsynchronized step.
This is the engineering treatment of the training-shaped side of the fork introduced in Chapter 1.1. We take the four defining properties of the workload — it is synchronous, dense, collective-dominated, and checkpointable — and derive the building from them: the parallelism regime and the collectives it generates; the scale-up domain and scale-out fabric that must carry them; the power density and supported cooling envelope that the selected rack requires; the reliability philosophy that checkpointing makes rational; the storage system that feeds the GPUs and absorbs the checkpoints; and finally the multi-datacenter and gigawatt-campus regime where a single run outgrows a single building. Accept the cluster only against a job-completion deadline that includes checkpoint writes, lost progress, restore, data input and collective stalls; a peak-rate test alone does not close it.
Pre-training as one tightly-coupled supercomputer
Training a frontier model is a single optimization loop run across an enormous machine. The model and its data are split four ways at once. Data parallelism replicates the model and shards the batch; every replica computes gradients on its slice and the replicas must agree before the next step — an all-reduce of the full gradient on every iteration. Tensor parallelism splits individual matrix multiplications across GPUs within a tightly-coupled group, generating all-reduce / all-gather traffic on the critical path of every layer. Pipeline parallelism splits the layers into stages across nodes, passing activations forward and gradients back, and lives or dies on how small the pipeline bubble stays. Expert parallelism, for mixture-of-experts models, routes each token to a subset of experts that live on different GPUs — an all-to-all shuffle that is now one of the heaviest collectives in modern training. The combination is called 4D (or 5D, adding context/sequence parallelism) parallelism, and it is why a pre-training cluster behaves as one organism rather than many.
One consequence governs the whole design: the job is synchronous, so it advances at the speed of the slowest participant. A single GPU that runs 10% slow — a thermal throttle, a flaky NVLink, a degraded optic — can slow the entire cluster’s affected critical-path phase by 10%, because every other GPU waits at the next collective barrier for the straggler to arrive. The whole-step penalty depends on that phase’s share and exposed overlap; it is not automatically a 10% loss over the entire run. This is the straggler tax, and it is the reason training facilities are engineered for uniformity and tail control rather than average performance. It is also why a single hardware failure does not degrade the job gracefully; it halts it, forcing a restart from the last checkpoint — unless the training implementation carries elastic membership or checkpointless recovery, which buys graceful degradation with replicated state and spare capacity.
Scale-up domain design and the scale-out fabric
AI clusters have two networks, and conflating them is the most expensive networking mistake in training. The scale-up fabric (NVLink and its NVSwitch fabric inside a node or rack) is the memory-coherent, ultra-high-bandwidth domain where GPUs talk as if they shared one address space. The scale-out fabric (InfiniBand or RoCEv2 Ethernet across the back-end) connects those domains into the full cluster. The per-GPU bandwidth gap between them is roughly an order of magnitude: NVLink5 delivers 1.8 TB/s per GPU bidirectional, or ~900 GB/s each way (a 72-GPU NVL72 rack aggregates ~130 TB/s), versus a ~400 Gb/s scale-out NIC at ~50 GB/s each way — about an 18x difference per direction. (Divide the 1.8 TB/s bidirectional figure by the NIC's one-way line rate and you get ~36x; the like-for-like per-direction gap is 18x.) The first principle of training-fabric design follows directly: keep the heaviest collectives inside the scale-up domain, because every byte you push onto scale-out is an order of magnitude more expensive in bandwidth and latency.
Scale-up domain size is therefore a first-class workload decision. An 8-GPU HGX node, a 72-GPU NVL72 rack, and the coming 576-GPU Rubin Ultra NVL576 domain across eight MGX racks are not just bigger boxes — each enlargement raises the ceiling on how much tensor and expert parallelism you can fit before spilling onto scale-out. A larger NVLink domain lets you fit a whole tensor-parallel group, or a wider set of MoE experts, inside the cheap fabric, which is exactly why the industry is racing domain size upward. The downside is blast radius and packaging: a 72-GPU rack is ~3,000 lb of wet hardware carrying 5,184 in-rack copper NVLink cables, and a single NVSwitch fault now degrades 72 GPUs instead of 8.
Published scale-out fabrics span 1:1, 2:1–3:1, and reported 7:1 examples; treat those as observations, not training/inference defaults. Chapter 8.5 derives the per-tier ratio through its three tests — locality, exposure and failure headroom, all read off the measured traffic matrix — and validates it on the target fabric. Synchronous collectives often justify high bisection; local inference often permits upper-tier oversubscription, but distributed MoE inference, KV movement, and prefill/decode disaggregation can demand more. A named 1:1-versus-2:1 source model estimates ~31% back-end cost difference; it is not permission to choose either branch without workload evidence. → Chapter 8.5 (topology, sizing, oversubscription); Chapter 8.2 (scale-up fabric); Chapter 8.4 (InfiniBand vs RoCE).
| Scale-up domain | GPUs / domain | Intra-domain bandwidth | Parallelism it unlocks | System power | Cost / blast radius |
|---|---|---|---|---|---|
| HGX node (8-GPU) | 8 | NVLink5 ~14.4 TB/s aggregate / node | TP up to 8; EP narrow | 14.5 kW consumption / 15 kW system max (NVIDIA DGX B300) | Enterprise and long-tail serving tier, not a frontier training path; air or DLC per selected OEM system; small blast radius |
| NVL72 rack | 72 | ~130 TB/s rack aggregate | TP + wide EP inside one rack | 132 kW nominal HPE GB200; Lenovo GB300 135 kW TDP / 155 kW peak; NVIDIA RA facility basis up to 142 kW | DLC mandatory; 72-GPU fault domain |
| Vera Rubin NVL72 rack (VR200) | 72 | NVLink 6, rack-scale | TP + wide EP inside one rack | 188 kW Max Q / 228 kW Max P (Pegatron); 330 kW cabinet facility design basis (NVIDIA DSX) | DLC only; 72-GPU fault domain |
| Vera Rubin Ultra NVL576 / Kyber NVL144 rack and NVL1152 8-rack system (roadmap) | 576 across eight 72-GPU MGX racks; 144 per Kyber rack; 1,152 across eight Kyber racks | NVLink 6/7, rack-scale + optical multi-rack | TP + EP + more DP inside scale-up | ~600 kW per Kyber rack (planning point, 800 VDC); ~4.8 MW per NVL1152 eight-rack system | 800 VDC + DLC; very large blast radius |
Power density and the supported cooling envelope
Density in a training hall is set by the accelerator generation and the scale-up domain you chose, and it lands you on one side of a discontinuity. An HPE GB200 NVL72 has a 132 kW nominal rack TDP, NVIDIA's GB300 reference facility design basis is up to 142 kW, and the Rubin Ultra / Kyber facility planning point is ~600 kW — qualify each against its published facility envelope before it enters the design basis. For the HPE GB200 NVL72, roughly 115 kW goes to liquid and 17 kW to residual air; that product record—not a universal rack-kW threshold—makes direct-to-chip liquid necessary. Choosing this dense rack therefore chooses liquid cooling — and it does so before you order a single GPU, because the slab loading, the selected facility-side heat-rejection path, and the pipe-rack space all have to exist first. → Chapter 5.1 (the density wall); Chapter 5.4 (DLC, the 2026 default).
That same heat split, together with the rack's airflow and inlet limits, also fixes the water side. The rack's maximum liquid inlet and maximum liquid return are separate acceptance limits, not an operating pair, and flow follows from the approved fluid, the liquid load and the design ΔT you declare rather than from either maximum. Chapter 5.1 owns that heat balance and Chapter 5.7 the fluid envelope, including the ~165–236 L/min worked band for this rack's 115 kW liquid share. A rack outside its supported thermal envelope can throttle, so correlate facility and IT telemetry against the declared operating point. → Chapter 5.7 (warm-water loops); Chapter 5.6 (CDUs and the secondary loop).
Reliability philosophy: checkpoint-and-resume, MTBF, and straggler economics
Reliability posture is where a training facility most needs two distinct models. Facility availability measures whether power, cooling, and connectivity remain inside their contracted states. Training goodput measures how much accelerator time advances the job after node failures, detection, checkpoint overhead, restart, and replay. Checkpoint-and-resume lowers the consequence of some interruptions, but does not make facility continuity irrelevant or select a topology. The design must model both.
The failure rates are not theoretical. Meta's Llama 3 405B run logged 419 unplanned interruptions over 54 days on 16,384 H100s — roughly one every three hours — with approximately 78% hardware-caused; the paper reports 58.7% as GPU issues, although its table's counts and printed percentages do not reconcile. SemiAnalysis reported roughly seven days of MTBF for one 512-H100 cluster at a top-tier operator in October 2024. That observation is not a per-GPU rate or scaling law; larger jobs expose more failure domains, but their effective interruption distribution must be measured for the named fleet, job, software stack, and event definition. Alibaba's production study put the large-job failure rate near 43% with ~73% recoverable via restart. At these rates, the cluster is always healing — so the design question is not 'how do we prevent failures' (you cannot) but 'how do we make each failure cheap.'
That reframing changes the economics of the redundancy decision without making it automatic. Compare each defined maintenance and fault state, transfer interruption, post-event loading, path and control independence, common modes, and recovery time against the service objective. Checkpointing, hot spares, and fast detection can make N, N+1, or distributed power economic for a named design, while a contract or recovery limit can still require uninterrupted paths. The anti-pattern is selecting either 2N or a lean topology from “checkpointable training” alone. → Chapter 12.2; Chapter 12.5.
The straggler economics close the loop. Because the job runs at the slowest GPU's pace, a partially-degraded node is often worse than a dead one — a dead node is evicted and replaced, but a silently-slow node taxes every step until it is detected. Mature operators therefore invest heavily in tail telemetry: per-GPU thermal and clock monitoring, NVLink and optic error counters, and collective-timing instrumentation that flags the straggler before it has bled hours of goodput. The facility's job is to give that telemetry nothing to find — uniform cooling, uniform power, no thermal hot spots — because every degree of thermal non-uniformity across the hall is a latent straggler. → Chapter 14.2 (DCIM and telemetry); Chapter 10.6 (GPU health observability).
| Decision axis | Training (checkpointable) | Inference (always-on) | Why they diverge |
|---|---|---|---|
| Primary objective | Goodput (useful FLOP-time) | Availability vs latency SLO | Job restarts vs lost revenue |
| Facility power | State-based selection: maintenance/fault continuity, interruption and recovery SLO | State-based selection: maintenance/fault continuity, interruption and fleet failover SLO | Checkpointing and request failover change outage consequence; neither selects topology |
| Back-end fabric | Measured collective traffic + step-time/MFU target | Measured request/KV/EP traffic + tail-latency SLO | Coupling, placement, and failure headroom set bisection—not the label |
| Failure response | Evict, hot-spare, resume from checkpoint | Route new requests; recover or retry in-flight work within the service budget | Synchronous halt vs independent requests |
| Where extra $ goes | Faster checkpointing, spares, straggler detect | Redundant power/cooling, geo-distribution | Goodput nines vs availability nines |
Checkpointing as a training constraint (its bearing here)
The full optimal-interval mathematics — the Young/Daly result that sets the checkpoint cadence balancing checkpoint cost against expected lost work — is canonical and lives in Chapter 9.4. Here we cover only its bearing on the synchronous training building, which is twofold and concrete.
First, checkpointing changes the outage consequence and can strengthen the economic case for N or N+1, but it does not select the facility topology. The state-based model must still prove which maintenance and fault cases may interrupt load, for how long, at what post-event capacity, and with what recovery and common-mode exposure. The cheaper and faster the checkpoint, the less work a failure costs, the more aggressive the cadence you can afford, and the lower the goodput penalty of any given MTBF. That makes checkpoint bandwidth a first-class facility requirement: the storage system must absorb a full-cluster checkpoint — terabytes of optimizer and model state — fast enough that the GPUs stall only briefly, because every GPU is idle during a synchronous checkpoint barrier. A slow checkpoint path quietly converts into lost goodput on every interval. → see storage, below, and Chapter 9.3 (GPUDirect Storage).
Second, the cadence interacts with the failure rate to size everything else. At one interruption every three hours (Llama 3 scale), a checkpoint cadence and a lost-work budget together determine how many hot spares you must keep warm and how fast the orchestration plane must detect, evict, and re-place a failed node to keep measured useful progress inside the job-completion budget. The facility decision that flows from this: provision the checkpoint storage tier and the spare-node pool as deliberately as you provision GPUs — they are the levers that convert a high failure rate into high goodput. → Chapter 10.7 (autonomous recovery); Chapter 14.6 (spares strategy).
Storage for training: checkpoint bandwidth, dataset streaming, and LOSF
A training cluster's storage exists to keep expensive GPUs fed and to absorb checkpoints without stalling them — and it has three distinct jobs with different performance shapes. Checkpoint write bandwidth is bursty and enormous: at a synchronous barrier the whole cluster writes its state at once, so the storage must sink terabytes in seconds to minimize the idle window. Dataset streaming (the data-loader path) is sustained, read-heavy, and latency-sensitive in the tail — if the loader cannot keep every GPU's input queue full, the GPUs starve and MFU drops, the same straggler logic applied to data instead of compute. Object/capacity storage holds the raw corpus and cold checkpoints. The parallel file system (Lustre, GPFS/Storage Scale, WEKA, VAST and kin) sits in front, and increasingly the data-loader and checkpoint paths use GPUDirect Storage to move bytes directly into GPU memory, bypassing the CPU bounce buffer. → Chapter 9.1 (why storage determines GPU efficiency); Chapter 9.2 (parallel file systems); Chapter 9.5 (the data-loader path).
LOSF — Lots Of Small Files — is the storage pathology specific to AI training, and it is worth naming because it ambushes teams that sized for bandwidth alone. Training corpora and tokenized shards are frequently millions of small objects, and small-file workloads are bound by metadata operations and IOPS, not by sequential throughput. A file system tuned for the big sequential reads of checkpoint restore can choke on the random small-file reads of dataset streaming, leaving GPUs starved while the bandwidth meter reads low. The facility consequence: the storage tier must be specified against the LOSF and checkpoint-burst profiles explicitly, not against an average GB/s number — and the metadata path (often the silent bottleneck) sized as deliberately as the data path. → Chapter 9.8 (sizing and data gravity); Chapter 9.9 (offline data-prep).
Deep dive: why the checkpoint-storage tier is a goodput lever, not a cost center
Treat checkpoint storage as commodity capacity and you leave goodput on the floor at every interval. The chain is mechanical. A synchronous checkpoint is a stop-the-world event: every GPU in the run holds at a barrier while cluster state is flushed, so the wall-clock cost of a checkpoint is (state size ÷ effective write bandwidth) multiplied across the whole fleet's idle time. Halve the write bandwidth and you double the idle window on every checkpoint — and you must checkpoint frequently because the failure rate is high. Under the uniform failure-in-interval assumption used in Chapter 9.4, expected replay is half the interval; so a slow checkpoint path forces a longer interval to amortize the stall, which in turn raises the expected lost work per failure. Slow storage thus costs goodput twice: once in the stall and again in the larger rollback.
The mitigations are all facility-and-stack decisions made at scoping time. Asynchronous / in-memory checkpointing stages state to host memory or NVMe and flushes in the background so GPUs resume almost immediately. Hierarchical checkpointing writes frequent local checkpoints (to node NVMe) and infrequent global ones (to the parallel file system), bounding both stall time and blast radius. GPUDirect Storage removes the CPU bounce buffer from the write path. Each of these trades a little complexity or local-NVMe capacity for goodput — and on a cluster where a point of goodput on a tens-of-thousands-of-GPU run is worth millions in GPU-hours, price the avoided lost progress against memory, NVMe, fabric and recovery complexity. The number to carry: provision checkpoint write bandwidth against the stop-the-world stall budget you can tolerate, not against steady-state throughput. → Chapter 9.4 (Young/Daly cadence math); Chapter 9.3 (GPUDirect). The acceptance test must also prove where the staged checkpoint survives and how long restore takes; a fast local acknowledgment is not a durable cross-domain checkpoint.
Scope & caveats
NVIDIA's published figure (GTC 2025) is 600 kW per Rubin Ultra Kyber rack and GTC 2026 did not revise it. SemiAnalysis (2026-05-26) reports Kyber Ultra 'approaching 660 kW' — a single-source analyst estimate for a 2027 part, recorded here rather than adopted, since the vendor primary figure still stands.
Scope & caveats
Select on the named door/rack, air and water conditions, fan state, containment, heat-capture target, residual room heat, climate/rejection, serviceability, redundancy, and future density.
Reference capacity, not a universal ceiling; verify named door/rack, water and air conditions, fan state, containment, and capture target.
Scope & caveats
Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.
Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.
Scope & caveats
A named source comparison, not training and inference defaults. Derive each project tier from measured traffic, collective/request mix, placement and overlap, topology, failure headroom, and step-time or tail-latency SLO.
Scope & caveats
Nominal directional endpoint ratios: 900 GB/s divided by 50 GB/s for the specified GB200 profile; divided by 100 GB/s for the specified GB300 profile. Not collective or application speed ratios.
Scope & caveats
Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.
The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Multi-datacenter and gigawatt-campus training
A single building has a ceiling — on power it can energize, on land, on a coherent cooling plant — and frontier runs have already hit it. The binding constraint of 2026 is power, and a single campus increasingly cannot host enough of it for the largest runs, so the supercomputer is being stretched across multiple buildings and multiple campuses. Google has trained Gemini across multiple sites within and across data centers, connecting TPU superpods over its intra- and inter-cluster network with latency and bandwidth sufficient to preserve a synchronous training paradigm — model-parallel within a superpod, data-parallel across superpods — and explicitly cites resilience (a power event at one site does not kill the run) and the simple physics of space and power as the reasons to distribute. Gigawatt-scale campuses (the multi-site Ohio build summing toward ~1 GW) are the current expression of this. → Chapter 8.8 (scale-across: multi-campus and cross-region fabric).
This forces a real fork. Synchronous cross-DC training keeps the single-job, lock-step model and extends the non-blocking fabric across campuses with dedicated dark fiber — preserving model quality and simplicity, at the cost of needing enormous, low-latency, high-bandwidth inter-campus links and tolerating the speed-of-light latency floor between sites (which bounds how far apart they can sit before the collective stalls). Asynchronous / low-communication training (the DiLoCo family — Streaming DiLoCo, DiLoCoX, and async variants) lets each site take many local steps before exchanging compressed pseudo-gradients, slashing inter-site bandwidth by orders of magnitude (DiLoCo demonstrated comparable quality while communicating ~500x less) and tolerating wide-area links and stragglers — at the cost of algorithmic complexity, staleness management, and a model-quality regime that is still maturing. The choice rewrites the inter-campus fiber plan, the power-resilience design, and the orchestration plane. → Chapter 8.8; Chapter 10.8 (training frameworks).
Deep dive: the power-resilience case for distributing a single run
Distributing a training run across campuses is usually read as a capacity story — no single site has the megawatts. But there is a second, subtler driver that matters as runs reach gigawatt scale: power resilience. A synchronous job on a single campus is hostage to that campus's power: a grid fault, a generator trip, or a ride-through failure can drop the entire run, and large data-center loads have demonstrably caused multi-hundred-megawatt instantaneous loss events that stress the grid (the July 2024 Virginia event: ~1.5 GW of data-center load dropped during a six-fault, 82-second reclosing sequence — the case behind NERC's rare May 2026 Level 3 alert). Spreading the run across sites degrades rather than kills the job only where the implementation earns it: state replicated so no campus holds an indispensable shard, elastic membership and quorum rules that let the collective re-form without it, and a recovery path you have tested. Meet those and the surviving campuses continue while the affected campus rejoins; miss them and distribution has added a failure domain rather than removed one, because asynchrony tolerates stragglers and staleness but does not rebuild lost state. Score the option on recovery time and remaining throughput.
The consequence for the building program: at gigawatt scale, the multi-datacenter decision is no longer purely about fitting the load — it is also a reliability-engineering decision that trades inter-campus fiber and orchestration complexity for independence from any single point of grid failure. That reframing pulls the energy-supply and ride-through strategy (→ Chapter 3.4, Chapter 4.10) and the cross-region fabric (→ Chapter 8.8) into the same design conversation as the parallelism strategy. The largest runs of 2026 are being scoped by people who hold all three at once.
Anti-patterns specific to training builds
The recurring training mis-scopes all share a root: reasoning from the equipment or the building instead of from the synchronous-dense-checkpointable nature of the job. Four are worth naming.
- Unvalidated back-end blocking. Copying 1:1, 2:1, or 3:1 from a workload label instead of measuring collectives, placement, overlap, failure headroom, and the step-time target. Either overbuilds stranded bisection or starves the job. Validate every tier on the target fabric. → Chapter 8.5.
- Topology chosen from checkpointability. Treating “checkpointable” as proof that N/N+1 is enough—or as proof that 2N is waste—without modeling maintenance states, defined faults, transfer interruption, post-event loading, independence, common modes, recovery SLO, and contract. Checkpointing is one consequence-mitigation input. → Chapter 12.2.
- Designing to today's density. Pouring a slab and water plant for the current generation, then being unable to absorb the next density step without re-pouring concrete. Price the irreversible reserve for each named ramp scenario and record when it is released. → Chapter 5.10.
- Bandwidth-only storage spec. Sizing the storage tier against an average GB/s number while ignoring the LOSF metadata profile and the stop-the-world checkpoint burst — starving GPUs on the data-loader path or stalling them on every checkpoint. → Chapter 9.1.
Cite this chapter
Fehn, J. (2026). Training Data Centers: Synchronous, Dense, Checkpointable (Chapter 1.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-1-strategy-workload-archetypes-and-economics/1-2-training-data-centers-synchronous-dense-checkpointable (accessed 2026-09-29).
@misc{aidc-1-2,
author = {Fehn, Jacob},
title = {Training Data Centers: Synchronous, Dense, Checkpointable (Chapter 1.2)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-1-strategy-workload-archetypes-and-economics/1-2-training-data-centers-synchronous-dense-checkpointable},
note = {Accessed 2026-09-29}
}