The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 8.8

In this chapter · 6 sections
Term help

Scale-Across: Multi-Campus & Cross-Region DCI for Portfolio and Qualified Cross-Site Workloads

Scale-across is an optional portfolio-connectivity tier; decide whether sites need DCI and, only for eligible work, whether a job may cross the latency, bandwidth, algorithm, placement, and correlated-failure boundary.

POWER-BOUNDGOODPUT

What you'll decide here

  1. Whether the portfolio needs multi-site DCI and for which traffic; if a job will cross sites, declare its algorithm, latency and failure tolerance, placement plan, and radius before selecting transport.
  2. Synchronous scale-across (treat remote campuses as one fabric and pay the bandwidth and exposed-latency tax) versus DiLoCo/local-SGD/hierarchical algorithms that exchange less often — compare actual inter-site bytes and dependent rounds against surviving payload capacity, and buy the traffic saving only after optimizer changes meet the same model-quality target.
  3. Which DCI optics and transport carry inter-site traffic: a named grey 800G client PMD over dark fiber, 800ZR or an interoperable ZR+ coherent mode over DWDM, or full OTN/managed transport — each carries a different fiber bill and operational owner. Qualify host/FEC compatibility, optical margin and restoration on the actual protected route.
  4. Where the fault domains and checkpoint boundaries sit once a job spans regions: a fiber cut, a regional power event, or a site evacuation is now a correlated failure that flat intra-DC checkpointing did not plan for.
  5. Whether the second site is a true training peer (it joins the synchronous/relaxed-sync run) or a checkpoint-and-burst replica (it holds state and absorbs overflow), because that decision sets the inter-site bandwidth you must buy and light.

Two chapters of this Part built the network inward-out: the scale-up domain that binds a named accelerator group, such as GB200 NVL72’s 72 GPUs, through supported peer operations and collectives (Chapter 8.2), and the scale-out Clos or other selected fabric that stitches scale-up domains into the campus’s larger GPU cluster (Chapter 8.5). This chapter is the third tier — scale-across — and it provides optional DCI among sites for portfolio capacity, replication, service failover, burst or overflow capacity, and qualified cross-site computation. The binding constraint of the 2026 era is megawatts at the substation, not accelerators in the supply chain (Chapter 16.1). Portfolio capacity can split among sites without forcing one synchronous job to span them. A job crosses only when its algorithm, latency and bandwidth budget, placement plan, and correlated-fault model are explicitly qualified for that boundary.

Cross-site training is one demanding use case, not the definition of the tier. Google trained Gemini Ultra across multiple datacenters on TPUv4 pods, combining SuperPods over its intra- and inter-cluster network with latency and bandwidth sufficient to keep the run synchronous; that is the origin case, and the hardware under the tier has turned over since. The campus-scale mix is not all NVIDIA: TPU, Trainium, AMD Helios, and other ASIC systems deploy at the same scale, and TrendForce puts ASICs at 27.8% of 2026 AI-server shipments (2026-03-18). OpenAI's Stargate program is building gigawatt-class capacity across multiple U.S. sites. If a job will cross sites, the latency floor, optics, transport, algorithm, placement, and correlated-failure model become one qualified design; otherwise DCI can serve portfolio traffic without joining the sites into one synchronous machine.

The latency/bandwidth hierarchy: scale-up >> scale-out >> scale-across

The three-tier hierarchy is the core mental model for this whole Part, and scale-across is the tier where the numbers drop by orders of magnitude. Inside scale-up, NVLink supplies the platform-specific peer operations and per-GPU bandwidth recorded in Chapter 8.2. Scale-out hands exchanges to the selected 400–800 Gb/s NIC paths, with topology, software and workload costs budgeted in Chapter 8.1; neither a switch-hop latency nor a peak interface rate is an application collective time. Scale-across drops again, by orders of magnitude on both axes at once: per-site uplink bandwidth is whatever DWDM you can light and pay for (tens of Tb/s aggregate is a large buy), and the latency floor is set by physics you cannot engineer around — ~5 µs per kilometer of fiber, one way.

That 5 µs/km figure governs everything downstream in this chapter. Light moves at ~300 m/µs in vacuum but ~30% slower in silica (refractive index ~1.47), so standard SMF-28 fiber adds ~4.9 µs/km one-way — call it 5. A campus-to-campus link 40 km apart has a ~200 µs one-way floor, ~400 µs round trip, before a single transponder, amplifier, FEC block, or switch hop adds its own delay. Two regions 1,000 fiber-km apart sit at ~5 ms one-way, ~10 ms RTT. An all-reduce that fits a short local exchange can spend milliseconds waiting on a WAN route, and a synchronous step pays that delay for each dependent round left on its critical path. MFU falls when those waits displace useful compute; large local compute intervals can absorb delays that a fine-grained collective cannot. The entire algorithmic apparatus of this chapter — relaxed synchrony, gradient compression, hierarchy — exists to hide or amortize that floor.

Scale-up, scale-out and scale-across, made spatial: local accelerator domains, a campus fabric, then protected fiber routes between sites. Illustrative — stated assumptions. Reservation widths are schematic, not rates. Recompute exposed byte-transfer time and dependent-round delay after a named route loss, with training and checkpoint traffic allocated separately. Verify shared duct, power and termination dependencies. Keep a committed restorable state; a healthy aggregate line rate alone does not admit the job. Exposed communication is max(zero, bytes / reserved payload + dependent rounds × RTT − valid overlap); add other work and compare with the declared deadline.

The first fork: do you go multi-site at all?

Before any optics or algorithm, decide why the portfolio needs multiple sites. Separate independent jobs, replication, service failover and overflow capacity from the decision to split one training job. A power or delivery constraint can make another site valuable without justifying a synchronous WAN dependency. If a job must cross, price the exposed communication and correlated-failure cost against the alternative schedule, using the project’s utility and equipment need-by dates. It is still a race between two clocks: depreciation on accelerators already bought and the utility-service/upgrade schedule for the megawatts they need. Idle silicon loses value while it waits, but splitting one job adds exposed communication and correlated-failure cost; independent jobs at the second campus may capture its power without that WAN dependency. → portfolio rationale in Chapter 16.1; siting and delivery in Chapter 3.2.

Once you've decided to split, the next question is what the second site is for, and the two answers buy very different amounts of fiber. A training peer site runs part of the synchronous (or relaxed-sync) job and must exchange gradients or pseudo-gradients with the primary on a cadence the algorithm sets — bandwidth-hungry and latency-sensitive. A checkpoint-and-burst replica holds a copy of model state and absorbs elastic inference or batch overflow — it needs enough bandwidth to ship checkpoints (gigabytes to terabytes on a slow cadence) but is indifferent to per-step latency. Mis-classifying the relationship is a classic over-spend: peer-grade DWDM for a site that only ever needed checkpoint-grade bandwidth, or the reverse — starving a true training peer and watching the run stall on a link you under-bought.

DCI admission by service and protected-route evidence
ServiceBudget inputsOptical boundaryTransport contractAdmission evidence
Independent-site services / replicationTransfer and user deadline, not a collective RTT targetQualified grey client optics or coherent lineOwned fiber, leased wave or managed serviceReserve payload and prove restorable state; Chapter 9.4
Synchronous cross-site jobBytes per step + dependent RTTs − valid overlapNamed PMD and host/FEC mode; line budget separatelyProtected capacity after shared-risk lossPass the step deadline in both healthy and failed-route states
Relaxed / hierarchical trainingExchange volume, interval, burst deadline and quality targetQualified line mode on each surviving routeCapacity reservation plus restart/rejoin procedurePass optimizer-quality and state-consistency tests, not distance alone
Service failover / burstDetection, state availability and redirected demandClient optics, coherent equipment and protection modeActual route and service owner; test feedback pathVerify access, keys, restoration and survivor capacity before admission
Distance contributes propagation; payload reservations, dependent rounds, route recovery and application quality determine admission. Confirm physical diversity, including entrances, ducts, amplifiers and power.

Coherent DCI optics: how the inter-site traffic actually moves

Inside the hall, the interconnect debate is copper-versus-optics over single-digit meters (Chapter 8.9). Across sites, traffic travels kilometers of glass, and the debate is which qualified grey or coherent optical application carries it together with the required protection and transport service. The fork has three branches, and they trade reach, cost-per-bit, and how much transport plumbing you own.

Grey optics over dark fiber light a point-to-point application without wavelength multiplexing in a line system. Pick the exact PMD and reach: a four-lane parallel DR4 application uses four transmit/receive pairs, whereas a four-wavelength FR4/LR4 application combines wavelengths onto one pair; each rate needs its own PMD definition. Verify that distinction before pricing pairs or patching. Coherent ZR/ZR+ pluggables move line termination into a supported router or switch port, reducing separate transport equipment when the host supplies the power, cooling and control interface. OIF’s October 2024 800ZR agreement targets single-span amplified 80–120 km DWDM; fitting coherent line termination in a supported QSFP-DD or OSFP port can replace separate transponders, but that target still needs the line budget. A ZR+ label alone does not establish a regional reach or a shared interoperability mode. Full OTN / managed transport adds a separate service boundary for grooming, protection and operations; choose it when those responsibilities belong with the transport provider. → channel physics in Chapter 8.9; fibers and repair access in Chapter 8.10.

Deep dive: why 800ZR collapsed the DCI cost structure — and what it doesn't fix

The economic change from coherent pluggables is an ownership change as much as a packaging change. Putting the coherent DSP and optical engine in a supported QSFP-DD or OSFP module lets a compatible router port terminate the line directly — IP-over-DWDM — avoiding a separate transponder chassis and its rack space, power and operational handoff. That saving must include supported host power class, cooling, software, diagnostics and spares; moving the module onto a faceplate does not remove the line-system work. The market forecast in the tile below is a dated supplier-market scenario, not a design operand or a guarantee of cheaper delivered service.

What coherent pluggables do not remove is propagation: an illustrative 120 km fiber route adds about 600 µs one way at 5 µs/km, before FEC, DSP, switching and queueing. Lighting more 800G wavelengths adds bandwidth, not a faster clock for each dependent round. Identify the Ethernet client attachment, client mapping and FEC, then the coherent line mode, wavelengths, amplifiers, receive power and OSNR limits separately from passive insertion loss. Assign monitoring, protection switching, fault localization and restoration, including encrypted service and key availability from Chapter 11.8. Test both directions under load and on the protection route. ZR optics can make the pipe cheaper to build; they do not make the WAN disappear or supply the reserved payload a cross-site collective needs.

The central fork: synchronous scale-across vs relaxed-synchrony algorithms

The rest of the chapter's decisions follow from this one. Once the job spans a metro or regional radius, you must choose how synchronization happens across the thin WAN pipe, and the choice trades convergence quality and simplicity against inter-site bandwidth and tolerance to latency.

Branch A — synchronous scale-across. Treat the remote campuses as one extended fabric: qualified data- and pipeline-parallel groups may span sites while tensor- and expert-parallel groups stay site-local, and every step still ends in a global all-reduce that now traverses the WAN. This preserves the synchronous update structure; numerical reduction order and reproducibility still depend on the implementation, and it is what Google did for Gemini Ultra. The cost is acute sensitivity to inter-site bandwidth and latency: you must overlap communication with computation aggressively, place the parallelism dimensions so the chattiest collectives (tensor-parallel) stay inside a site and only the more tolerant ones (data-parallel, pipeline) cross the WAN, and accept that bisection bandwidth across sites is your hard throughput ceiling. Get the placement wrong and the run waits on WAN bandwidth or dependent round trips while MFU falls. This branch is viable wherever the protected-route calculation and measured overlap keep exposed exchange time inside the step budget.

Branch B — relaxed-synchrony (local-SGD family). Stop synchronizing every step. Each site runs many local optimizer steps independently, then the sites exchange and average their accumulated updates — "pseudo-gradients" — only periodically. This is the DiLoCo paradigm (Distributed Low-Communication): inner optimizer AdamW running H local steps, outer optimizer Nesterov momentum averaging across workers. DiLoCo on 8 workers matched fully-synchronous optimization on C4 while communicating ~500x less, and OpenDiLoCo reproduced this training across two continents and three countries at 90–95% compute utilization. Streaming DiLoCo (2025) overlaps the periodic communication with compute to cut peak bandwidth further; follow-on work (SparseLoCo) works out to roughly 30–94x less total traffic than DiLoCo and ~890–2,810x less than standard DDP, depending on sparsity. The cost is no longer bandwidth — it's convergence risk: relaxed synchrony introduces staleness and can perturb final model quality, the inner/outer optimizer and sync-interval are extra hyperparameters that must be tuned, and the failure modes are subtler than synchronous SGD's clean restart-from-checkpoint.

Branch C — hierarchical / async hybrids. A pragmatic middle to qualify between per-step WAN synchronization and independent local work is synchronous within a site, relaxed across sites. Each campus is one tightly-coupled synchronous fabric; the campuses themselves are coupled with local-SGD-style periodic averaging or bounded-staleness asynchrony. This matches the algorithm to the physics tier-by-tier — pay full synchronous cost only where bandwidth is cheap (intra-DC), and pay the relaxed-sync convergence tax only where bandwidth is expensive (cross-WAN). Gradient/pseudo-gradient compression (quantization, top-k sparsification, low-rank/momentum decoupling) layers on top of any branch to shrink what crosses the WAN further.

Synchronous scale-across vs relaxed-synchrony: the convergence-vs-bandwidth trade
ApproachInter-site trafficLatency toleranceConvergence quality riskAdmission condition
Synchronous scale-across (global all-reduce every step)Full exchange required by the selected collectiveExposed dependent rounds must fit the stepPreserves synchronous updates; verify numerical behaviorAny route pair that passes healthy and survivor budgets
Hierarchical (sync in-site, relaxed across sites)Periodic cross-site averagingDeclared exchange interval and burst deadlineValidate quality at the selected intervalRoutes whose survivor serves the qualified exchange schedule
DiLoCo / local-SGD (periodic pseudo-gradient averaging)~100–500x less than synchronous DDPHigh — H local steps between syncsModerate — staleness perturbs final qualityPass optimizer-quality and state-consistency tests at the selected sync interval, not distance alone
Local-SGD + compression (Streaming DiLoCo, SparseLoCo)~30–94x less than DiLoCo; ~890–2,810x vs DDPVery high — compute/comm overlappedModerate–higher; more hyperparameters to tuneSame quality and recovery tests at the compressed exchange schedule; the sparsity setting is part of the qualification
Traffic-reduction figures are reported research results (DiLoCo, OpenDiLoCo, Streaming DiLoCo / SparseLoCo, 2024–2025) and are workload- and config-dependent, not guarantees. "Quality risk" is relative to exact synchronous SGD on the same token budget.

If your sites are close enough that a global all-reduce fits the per-step compute budget, run synchronous scale-across: protect model quality, and pay for it in fat, expensive DWDM and aggressive overlap engineering. If they do not, buy more surviving payload, change placement, or qualify a hierarchical / DiLoCo-class algorithm with a less frequent or compressed exchange schedule. That trades fiber demand for convergence-tuning and recovery work; measure its traffic saving and validate the same model-quality target before accepting the trade. The naive failure is to light a thin WAN pipe and run an unmodified synchronous job across it, leaving both campuses' GPUs blocked on an all-reduce that physics won't let finish in time. → checkpoint and fault-domain consequences below.

Worked decision: admit DCI after one route loss

Healthy communication is 8.0 × 2³⁰ B / (40 × 10⁹ B/s) + 2 × 0.0020 s = about 220 ms. The step is 1,000 ms + max(0, communication − 150 ms), about 1.1 s: passes the model because the unrounded value is below 1,100 ms. After one route loss, communication is 8.0 × 2³⁰ / (20 × 10⁹) + 0.0040 s, about 430 ms; the step is about 1.3 s: fails. Compare the unrounded calculation to the deadline; the rounded healthy result does not mean equality.

Communication must fit 150 + (1,100 − 1,000) = 250 ms. Removing the 4.0 ms dependent-round cost leaves 246 ms for bytes, requiring 8.0 × 2³⁰ B / 0.246 s, about 35 GB/s of surviving training payload. Buy that survivor capacity or refuse synchronous admission. The synchronization flip streams 4.0 GiB per step at 20 GB/s: communication returns to about 220 ms and the step to about 1.1 s, again below 1,100 ms before rounding. This passes the timing model, but remains HOLD for algorithm-quality and recovery evidence. Merely sending a burst every second step would not prove either step’s deadline; the assumed streaming schedule must be demonstrated.

Preserve the reservation after failure. If an assumed complete 1.0 TiB checkpoint is produced every 900 s and receives 10 GB/s of the surviving reserved service, transfer takes 2⁴⁰ B / (10 × 10⁹ B/s), about 110 s, within the interval. That is a transfer check only; committed model, optimizer, RNG and metadata state plus a successful restore belong to Chapter 9.4. Verify route diversity by duct, entrance, line equipment and power, then measure failover and reverse feedback under load. The OIF client/line mapping supplies the transport boundary; Chapter 13.7 executes the installed acceptance.

~5 µs/km
one-way fiber propagation floor (SMF-28, ~4.9 µs/km); ~10 ms RTT per 1,000 km — the latency physics scale-across cannot engineer around
~500x
less inter-site communication for DiLoCo (8 workers) vs fully-synchronous SGD, matching convergence on C4
90–95%
compute utilization sustained by OpenDiLoCo training a model across two continents and three countries
~30–94x / ~890–2,810x
total traffic cut for SparseLoCo (H=15, 2-bit) vs 8-bit DiLoCo and 16-bit AdamW DDP — derived from Table 1 message sizes × sync counts; density-dependent (3.12% → 0.78%)
80–120 km
OIF 800ZR amplified single-span DWDM reach; qualify the client, line and protected route
Scope & caveats

Client mapping and coherent line-interface agreement. Reach and survivor payload must be engineered on the actual line system; no generic ZR+ extension is implied.

>200k units / >$1Bforecast
forecast 2026 shipments / revenue for 800G coherent optics; pluggable-coherent module market ~$2B (2025) → ~$5B (2029)
~7 GWforecast
planned Stargate capacity across the Abilene flagship + five new US sites — multi-campus by necessity, not choice

Cross-site fault domains: a new and correlated failure model

Intra-DC reliability engineering already carries correlated failure — the shared bus, the shared CDU, the fleet-wide firmware push, priced as a beta-factor in Chapter 12.5 — but its day-to-day arithmetic runs on the independent kind: a GPU dies, a NIC flaps, a cable goes bad, and you ride it out with hot spares and checkpoint-and-resume (Chapter 12.2). Scale-across adds shared-risk groups wider than any inside one building. A fiber cut between campuses, a regional grid event, a utility curtailment, a metro-wide cooling-water issue, or a site evacuation takes out an entire site's worth of accelerators at once — one correlated failure spanning thousands of GPUs. The DCI link itself becomes a first-class fault domain. Flat intra-DC reliability posture no longer covers you.

Three things change. Path diversity stops being optional: inter-site fiber must be physically diverse (separate conduits, separate entrances, ideally separate carriers) or the WAN link is a guaranteed correlated-failure single point. Checkpoint placement must become site-aware: a checkpoint that lives only on the campus that just lost power is no checkpoint at all. Relaxed synchrony leaves recent model copies at each site between exchanges, which can become a recovery advantage only with durable optimizer/RNG state, agreed membership and a committed global exchange boundary. Test restart from that boundary and rejoin after site loss; a local model copy alone is not a restorable global checkpoint. Blast-radius accounting must treat "site" as the failure unit: not "what happens when a node fails" but "what fraction of the run survives when a whole campus drops, and how fast can the remaining sites continue or resync." → the checkpointing math — interval, bandwidth, and the storage/network trade behind multi-region durability — is worked in Chapter 9.4; the goodput-vs-availability reframing in Chapter 12.2.

Putting it together: the scale-across decision sequence

The chapter resolves to an ordered sequence of forks, each constraining the next, now stretched across the WAN.

  • Does one job have to span sites? Portfolio capacity, replication, failover and independent jobs split across sites without paying any synchronous-training penalty — size their DCI on data movement, durability and user latency, and the algorithm fork below does not apply to them. A single job crosses only when you are power-bound on one campus and the silicon clock is beating the power clock; the goodput tax on that job is real and unrecoverable. → Chapter 16.1.
  • What radius can you energize? Intra-metro, metro, or regional — set by which substations you can interconnect, not by preference. The radius picks your latency floor at ~5 µs/km. → Chapter 3.2.
  • Which DCI optics and transport? Choose a named grey client PMD, coherent line mode and protection service whose host, optical and recovery limits close on the route. → Chapter 8.9, Chapter 8.10.
  • Synchronous or relaxed-synchrony? Synchronous only if surviving payload and dependent rounds leave the all-reduce inside the step budget; otherwise buy capacity or qualify a changed exchange schedule and model quality. This is the model-quality-vs-bandwidth fork. → algorithms above.
  • How do fault domains and checkpoints span sites? Route-diverse fiber, site-aware checkpoint placement, blast-radius measured per site. → Chapter 9.4, Chapter 12.2.

Walk the sequence and the multi-campus run becomes a chain of well-posed decisions, each with a named downstream cost. Skip the first fork — spanning one job across sites when only your capacity, not your job, needed to move — and you pay a permanent goodput tax for nothing. Skip the protected-route calculation and you can light expensive fiber only to watch both campuses idle on an all-reduce the surviving bandwidth and round-trip delay cannot finish before its deadline.

Scale-across is the third tier above the scale-up domain of Chapter 8.2 and the intra-DC scale-out Clos of Chapter 8.5; congestion control and collectives whose dependent delay this chapter budgets are engineered in Chapter 8.6, and the out-of-band/timing fabric in Chapter 8.7. The coherent-optics and transport primitives — modulation, FEC, link budgets, baud rates — are taxonomized in Chapter 8.9, with the fiber plant and structured cabling in Chapter 8.10. The checkpoint math behind cross-region durability lives in Chapter 9.4; the correlated-failure reliability rethink in Chapter 12.2. The power-bound rationale that forces multi-site in the first place is the macro story of Chapter 16.1, with siting and grid-queue mechanics in Chapter 3.2 and the optics roadmap consolidated in Chapter 16.2.
Cite this chapter
Fehn, J. (2026). Scale-Across: Multi-Campus & Cross-Region DCI for Portfolio and Qualified Cross-Site Workloads (Chapter 8.8). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-8-scale-across-multi-campus-and-cross-region-fabric-dci-for-distributed-traini (accessed 2026-09-29).
@misc{aidc-8-8,
  author       = {Fehn, Jacob},
  title        = {Scale-Across: Multi-Campus & Cross-Region DCI for Portfolio and Qualified Cross-Site Workloads (Chapter 8.8)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-8-scale-across-multi-campus-and-cross-region-fabric-dci-for-distributed-traini},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit