The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 8.5

In this chapter · 10 sections
Term help

Scale-Out Topology, Sizing & Oversubscription

Oversubscription and topology turn a GPU count into a bill of materials, a blocking factor, and an MFU ceiling; the 2026 error is setting one ratio for the whole fabric, or taking it from the workload's label instead of its measured locality, exposure and failure headroom.

GOODPUTPOWER-BOUND

What you'll decide here

  1. Which topology family — rail-optimized fat-tree, rail-only, single-tier, scheduled fabric (DSF/DES-class), multi-planar, dragonfly, or OCS/torus — your scale and your dominant collective actually demand, because the choice sets switch count, optics count, worst-case hop diameter, and bisection bandwidth all at once.
  2. The oversubscription ratio at every scale-out tier — start with the cited non-blocking training rail, then test a 2:1 spine or spare uplink capacity against locality, exposed communication and the required failure headroom. Remove switches and optics only where the actual port map and surviving service still meet the workload deadline; the saving is the equipment you can delete from that BOM.
  3. The scalable-unit (SU) boundary you build in — the repeatable GPU/switch/optic block you replicate — because sizing in SUs, not in ad-hoc GPU counts, is what keeps the BOM, the cabling, and the fault domains tractable as the cluster ramps.
  4. How your parallelism plan (TP/EP inside the scale-up domain, PP and DP across the scale-out fabric) maps onto the topology — get the mapping wrong and a non-blocking fabric still starves the collective it was built to serve.
  5. Where you spend the reach budget: copper inside the rack or SU where the qualified route fits, optics across longer spine routes — because at the selected 800G/1.6T application the electrical channel can constrain the floor plan before switch radix does. Count host/fabric channels, electrical and optical lanes, active fibers, patching and replaceable parts before pricing it.

By the time you reach this chapter you have already chosen a protocol and transport (Chapter 8.4) and the switch, NIC, and DPU silicon (Chapter 8.3). What remains turns those components into an actual machine: how do you wire N GPUs together, and how much bandwidth do you refuse to pay for? Those are two views of one decision. Topology fixes the shape — how many tiers, how many switches, how long the longest path, how much bisection bandwidth the cluster sustains when every GPU talks to every other at once. Oversubscription fixes the price — how much of that theoretical bisection you actually build versus deliberately starve, because the workload will never use it.

A topology choice propagates deterministically into a switch count, an optics count, a cable-length distribution, a power draw, and an effective bisection bandwidth that caps the collective throughput of the parallelism plan running on top of it. What follows builds the topology taxonomy, derives the sizing math in scalable units, works the per-tier oversubscription fork from the published starting points and the three tests that move a tier off them, and maps the parallelism strategy onto the wires. The back-end fabric is the second-largest line item after the accelerators, and a mis-sized decision here is paid quietly in lost goodput on every step of every job.

The topology taxonomy: the families and what each buys

A handful of fabric families account for essentially every AI cluster shipping in 2026. They trade the same three currencies against each other: bisection bandwidth (how much the fabric can move when the traffic is all-to-all), switch and optic count (the cost and the failure surface), and worst-case hop diameter (latency and the blast radius of congestion). No family wins on all three; the right one is a function of scale and of which collective dominates your workload.

Fat-tree / Clos is the default and the safe choice. A folded multi-tier Clos with the right radix delivers full bisection bandwidth at every tier — genuinely non-blocking — and scales by adding tiers: a 64-port switch (32 down, 32 up) builds a three-tier fat-tree to 65,536 endpoints (k³/4); NVIDIA's 144-port Quantum-X800 reaches 10,368 NICs in just two tiers, and tens of thousands of GPUs in three. (The Ethernet sizing datum moved in 2026 too: sixteen liquid-cooled 102.4 Tb/s SN6600-LD switches — 64×1.6T ports each — pack 1.64 Pb/s of switching into one 48U rack for Vera Rubin fabrics; CoreWeave, Jul 2026.) The cost is switch count and optics: a non-blocking three-tier Clos is the most switch- and transceiver-heavy topology there is, and at 100k-GPU scale the optics alone become a top-tier line item and the fabric's leading failure and flap source.

Rail-optimized fat-tree is the AI-specific refinement, and it is the dominant production pattern for NVIDIA-class clusters. Instead of one flat Clos, group corresponding GPU/NIC positions into rails — for example, eight rails for an assumed eight-GPU server with one attachment per GPU, with GPU i in each node assigned to rail i. A rail can span several leaves and spines: same-rail traffic stays one leaf hop only when both endpoints attach to that leaf. A DP ring can therefore be entirely rail-aligned while every edge crosses a spine. Cross-rail traffic can use a qualified endpoint-forwarding path, including NVLink/PXN where supported; it is not made free by the rail label. The saving comes from paths or capacity the workload no longer needs and the design actually removes. Rail-only pushes this further: it removes the spine tier entirely, keeping only the rail leaves. Cross-rail traffic does not disappear — it detours through the scale-up domain, the move NCCL's PXN path already makes on rail-optimized fabrics (hop the NVLink domain to a GPU on the destination's rail, then cross without changing rails — NVIDIA, 2022). The MIT CSAIL/Meta proposal that named the pattern cuts network cost 38–77% versus a full-bisection fat-tree (arXiv, 2023–24); the price is scope — it assumes the collective profile stays rail-aligned, so all-to-all-heavy or failure-rerouted traffic pays the scale-up forwarding overhead, and the design is a bet on workload stability, not a free lunch.

Dragonfly trades bisection for switch count. Groups of switches are fully connected internally and sparsely connected to each other, yielding at most three hops endpoint-to-endpoint with far fewer switches and optics than a fat-tree of equal size — but lower bisection bandwidth, which is why it favors HPC and latency-bound patterns over bandwidth-bound all-reduce. OCS and torus is the hyperscaler-custom lane: Google's TPU pods use optical circuit switches to reconfigure the physical topology on demand (a 3D torus or twisted variant), decoupling logical from physical wiring and cutting worst-case diameter — Google's reconfigurable approach reduces a 1,024-chip worst case from 16 hops on a plain 3D torus to ~7. OCS removes a whole tier of packet switches and their optics, but it is a vertically-integrated play most operators cannot buy off the shelf.

Two more families bracket the taxonomy from opposite ends. Single-tier is the degenerate case worth naming because it is the simplest correct fabric wherever it fits: the entire back-end is one switch — one modular chassis, or one fixed box per plane — and most of the congestion machinery of Chapter 8.6 becomes moot, because there are no external inter-switch links; egress incast, internal fabric limits and receiver-resource control still remain. The ceiling is chassis radix: 2026-generation modular AI spines reach 576×800G or 1,152×400G in one box (~460 Tb/s), so up to 576 accelerators at 800G — or 1,152 at 400G — connect directly to one chassis and skip the leaf-spine tier entirely, along with its optics and cabling (vendor datasheets, 2026). Scheduled fabrics (Meta's DSF; the productized DES/DDC class) extend single-switch semantics beyond one chassis: leaf and fabric elements built on deep-buffer VOQ silicon behave as one logical switch — ingress traffic is sprayed as cells across all fabric links, and egress credits are granted before a packet is admitted — so grant-based admission schedules transfers inside that fabric domain instead of relying only on reactive ECMP/DCQCN behavior; endpoint overload, receiver resources and the recovery tests of Chapter 8.6 still need owners. Meta reports its DSF at 18K×800G-GPU scale (Meta engineering, 2025), and productized scheduled fabrics list >27,000 800G ports in a single logical switch. The trade: the fabric's internals are single-vendor by construction, per-port cost is higher, and the scheduled domain is one administrative blast radius — the lock-in-for-determinism bargain the deep-buffer row of Chapter 8.3 prices.

Topology family → bisection, cost, diameter, fit
TopologyBisection BWSwitch / optic countWorst-case hop diameterBest fit
Fat-tree / Clos (non-blocking)Full (1:1) at every tierHighest — most switches & optics2 tiers ~3 hops; 3 tiers ~5 hopsGeneral training; the safe default at any scale
Rail-optimized fat-treeCapacity provisioned per rail; multi-leaf rail traffic crosses spinesSavings only for connectivity or capacity actually removedOne leaf if colocated; leaf–spine–leaf across leavesRail-aligned jobs whose actual cut and failure budgets pass
Rail-only (spineless)Provisioned on-rail; cross-rail requires supported forwardingLowest fat-tree variant — no spineOne hop only in a one-switch rail; cross-rail forwarding is separateWorkloads that never need cross-rail east-west
DragonflyLower than fat-tree; sparse global linksLow — far fewer switches/optics≤3 router hops on minimal routes; nonminimal routes differHPC, latency-bound, cost-sensitive at scale
OCS / torus (reconfigurable)High within a pod; topology-on-demandLowest packet-switch count; OCS replaces a tierReduced vs static torus (e.g. 16 → ~7)Vertically-integrated XPU pods (TPU-class)
Single-tier (one switch; or one per plane)Full — the switch is the fabricLowest — no fabric tier at all1 hopClusters that fit one chassis radix (576×800G-class in 2026)
Scheduled fabric (DSF / DES / DDC-class)Capacity bounded by grants, internal fabric and egress serviceModerate — fabric elements replace the spineOne logical switchQualify admission, egress overload and scheduler/fabric failures
Multi-planar (N independent fabrics)Full per plane; aggregate across N planesPer-plane radix multiplies; two tiers held longer2-tier (~3 hops) held to 100k+ endpointsScaling two tiers past one fabric's radix (→ planes section)
Synthesis of NVIDIA DGX SuperPOD network-fabric reference, Meta RoCE/DSF, Google TPU OCS, and SemiAnalysis topology analyses, 2025-2026. 'Bisection' is the all-to-all worst case the fabric sustains; 'hop diameter' is endpoint-to-endpoint worst case. Single-tier/scheduled-fabric rows from vendor datasheets and Meta engineering (2024–26); multi-planar scaling is derived in the planes section below.

The table reads left to right as choice into consequence, the same cascade as the workload archetype in Chapter 1.1. A fully provisioned fat-tree buys bisection capacity with additional switches and optical interfaces, each adding its failure and repair exposure. Dragonfly accepts lower bisection for far fewer transceivers — a rational trade only if your dominant collective is not bandwidth-bound. Because the three currencies move together, sizing the topology is where the cost is set.

Illustrative — stated assumptions. The drawing shows a repeated rail from the chapter’s complete port ledger: equal-speed host attachments and uplinks, divided equally between spines X and Y. Solid arrows trace the cross-SU exchange through X; dotted paths and the cross show Y unavailable, with no junction at the bridged crossing. The step labels compare healthy service with the converged lost-spine case, under the chapter’s separate deadlines; transient recovery still needs its acceptance test. The switch and optical-module totals cover the complete repeated-rail build, including host module ends. Rail optimization requires participant mapping to these paths; cross-rail traffic needs a fully counted general fabric. A leaf loss disconnects its ranks and requires repair or relocation and restart.

Planes: parallel fabrics instead of deeper ones

A plane is a complete, independent copy of the scale-out fabric, and every accelerator connects to all of them. Split a NIC's bandwidth across its supported port configurations — an 800G NIC presenting as 2×400G, 4×200G, or 8×100G — and give each port to a different plane: N parallel leaf-spine fabrics, deliberately not interconnected, each carrying 1/N of every NIC's bandwidth, with the NIC or host stack balancing load across planes and containing per-plane failures. This is a deployed architectural pattern in the named reference designs; the selected implementation still has to meet its workload and recovery budgets. The B300-generation reference architecture attaches each GPU as 2×400GbE into two independent planes — on a NIC generation whose Ethernet ports top out at 400GbE, dual-plane is simply how 800 Gb/s attaches (→ Chapter 8.3) — with NCCL or the NIC's onboard plane load balancer spreading traffic, and, for two equal independent planes, a failed plane removing half the bandwidth instead of killing the job when endpoint recovery and the remaining capacity meet its deadline (NVIDIA reference architectures, 2025–26). The MRC/UET NIC generation extends the same move to four and eight planes with per-packet, hardware plane balancing (→ Chapter 8.4).

Keep the vocabulary straight, because the industry does not: a rail is a same-local-rank affinity group within one fabric — all the rank-i NIC ports across servers, mapped to as many leaves and spines as the endpoint population needs; one-hop reach requires the endpoints to share a leaf; a plane is which independent fabric a port belongs to. The two compose: each plane of a multi-planar cluster is itself rail-optimized, exactly as a single-fabric cluster would be.

The reason planes matter is arithmetic. In a symmetric non-blocking two-tier Clos — half of each leaf's ports facing endpoints, half facing spines — endpoint count scales with the square of switch radix: 64 ports of matching speed reach 2,048 endpoints; 512 reach 131,072. But radix at full port speed is capped by the switch ASIC: a 51.2 Tb/s chip is 64×800G or 512×100G, port count trading against port speed. Planes convert that trade into scale. Run each NIC at 1/N speed per plane and the per-plane radix multiplies: the same silicon that connects 2,048 endpoints as one 800G fabric connects 131,072 GPUs as eight 100G planes — the design published in the OpenAI-led MRC/SRv6 paper, which reports deployed four-plane (4×200G) and eight-plane (8×100G) systems, and the same two-tier-to-~128K arithmetic NVIDIA's multiplane analysis reaches (OpenAI et al., 'Resilient AI Supercomputer Networking using MRC and SRv6,' 2026; NVIDIA, 'High-speed Networking for Giga-Scale AI Factories,' 2026). Against a full-bisection three-tier build at that scale, the MRC/SRv6 paper counts roughly two-thirds the optics and three-fifths the switches, with a worst-case path of three switch hops instead of five to seven — which is why the frontier-scale consensus is hardening into two tiers, more planes, bigger spines, not a third tier. The price: per-plane speed drops (fine for spray-capable transports, a real constraint for single-path flows), the NIC must balance and fail over across planes, and the symmetry discipline of a fat-tree now applies N times over. → the congestion machinery a plane no longer shares in Chapter 8.6; the transports that made multi-plane native in Chapter 8.4.

Sizing in scalable units: the BOM math

The professional way to size a fabric is not to count GPUs — it is to count scalable units (SUs). An SU is the smallest repeatable block of GPUs, leaf switches, optics, and cabling that you replicate to grow the cluster. Sizing in SUs is what keeps the BOM, the cabling plan, the fault domains, and the procurement schedule tractable as a cluster ramps from one SU to dozens. In NVIDIA's GB200 SuperPOD reference, 1 SU = 8 rack-scale systems = 576 GPUs; a 16-SU cluster of 9,216 GPUs needs a switch BOM on the order of a thousand boxes across leaf, spine, and core for a non-blocking three-tier compute fabric, plus a separate storage fabric and a separate management fabric. The generic 512-GPU ledger below is a different platform and must not be used to order the GB200 reference.

The math that turns an SU into a BOM is mechanical once the radix and the blocking factor are fixed. For a non-blocking fat-tree built from radix-k switches, each tier must provide as much uplink bandwidth as downlink bandwidth — so for every port facing the GPUs there is a port facing the next tier up. That 1:1 uplink:downlink capacity rule is a necessary cut check for this symmetric leaf tier: half of each equal-rate leaf’s ports, and the optics on those uplinks, preserve upward capacity. It does not prove every traffic matrix is non-blocking; usable paths, distribution and endpoint limits still have to close, and top-tier spines have no higher tier to reserve ports for. Relax the rule to 2:1 (two downlinks per uplink) and you delete a large fraction of the spine switches and their transceivers — which is precisely the oversubscription lever the next section turns.

Three counting traps recur. First, the optics dominate the back-end BOM, not the switches — at 100k-GPU scale a cluster carries tens of thousands of miles of fiber and the transceiver line can rival the switch line. Second, the storage and management fabrics are not free: the compute fabric is the headline, but a production SU also carries a non-blocking storage fabric (e.g. 16x800G per SU) and an out-of-band management network (Chapter 8.7). Third, the reach budget, not the radix, often caps the physical topology: passive DAC dies around 1–2 m at 800G, active copper (AEC) stretches to about 7 m, and beyond that you are paying for optics on every link (Chapter 8.9 owns the reach ladder and its dates) — so the cable-length distribution implied by your floor plan is part of the sizing math, not an afterthought.

Deep dive: count one GB200 SU’s switches and optics on its own reference boundary

Take the NVIDIA GB200 reference and walk the count. One SU has eight NVL72 systems and 576 GPUs, with separate compute, storage and management attachments. For a non-blocking symmetric leaf tier, uplink and downlink capacity must match; half-down, half-up follows only for equal-rate ports. A hypothetical 64-port leaf with 32 GPU ports and 32 uplinks needs eighteen leaves for 576 endpoints, an ideal lower bound that omits OEM rail granularity. Replicate to sixteen SUs and that granularity sets the order: four groups of eight leaves and six spines require 32 leaves and 24 spines per SU, giving 512 leaves +384 spines +144 cores =1,040 compute switches. Storage and management remain separate. Ordering the ideal lower bound leaves this topology short; the generic case below uses its own port map.

Now the optics. Every spine or core link whose qualified route exceeds copper reach needs an optical channel with two module ends. At the NVIDIA reference's 9,216-GPU boundary, the transceiver bill therefore follows the actual optical port map; storage and management add their own ends. Drop to a 2:1 oversubscribed spine and recount exactly which spine switches, module ends and channels the port schedule lets you delete; the generic ledger below prices that hardware rather than assuming a one-third saving. The lesson: the SU is the counting unit and radix and blocking are the multipliers; the switch, optics and channel lines together decide what the back-end costs. → physical-layer reach and optic taxonomy in Chapter 8.9; structured cabling in Chapter 8.10.

Worked decision: a counted 512-GPU rail fabric

$20,000modeled
Illustrative switch unit price
Assumed per populated base switch, excluding optical modules.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching price/allowance, not a market quotation. Chapter 8.5 states installed/spare scopes and exclusions.

$800modeled
Illustrative DR4 module unit price
Assumed per module end.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching price/allowance, not a market quotation. Chapter 8.5 states installed/spare scopes and exclusions.

$400modeled
Illustrative installed channel or spare kit price
Defined passive scope; no module included.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching price/allowance, not a market quotation. Chapter 8.5 states installed/spare scopes and exclusions.

$3.0Mmodeled
Illustrative back-end capital allowance
Same exclusions as the ledger.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching price/allowance, not a market quotation. Chapter 8.5 states installed/spare scopes and exclusions.

Wire the complete port map before totaling it. For each rail r, leaf A ports 1–32 attach local GPU r in nodes 0–31; leaf B ports 1–32 attach that position in nodes 32–63. Leaf ports 33–48 go to spine X, using X ports 1–16 for A and 17–32 for B. Leaf ports 49–64 do the same on spine Y. Each cable endpoint is unique; X and Y ports 33–64 stay dark. There are 8×2=16 leaves and 8×2=16 spines. Each leaf presents 32×400 Gb/s down and the same up: 12.8 Tb/s per direction, 1:1. This is eight rail groups; the NIC is not independently attached to eight complete fabric planes.

512-GPU ledger: ports, optical channels and passive plant
ItemTraceInstalled quantity / scope
NIC ports64 nodes ×8512 host ports; NIC electronics excluded from price/power
Switches and ports16 leaves +16 spines; 64 cages each32 switches; 2,048 switch cages; 1,536 populated, 512 dark
Host channels64×8512 optical channels, 20 m each
Fabric channels16 leaves ×32 uplinks512 optical channels, 100 m each
Optical module ends2×(512+512)2,048: 512 host, 1,024 leaf and 512 spine
Lane / fiber map8 electrical lanes →4 optical Tx lanes/module8 active fibers/channel; unused MPO positions are not active strands
Active fiber length(512×20 +512×100) m ×861,440 channel-m; 491,520 strand-m =491.52 strand-km
Segments / panel couplings1,024×3 segments; 1,024×2 panel couplings3,072 segments; 2,048 panel couplings; 4,096 mated pairs including module ends
Cold spares2 switches; ceil(0.05×2,048); ceil(0.02×1,024)2 switches, 103 modules, 21 full channel kits; no operating power
The TIA FOTC DR4 application overview establishes four optical lanes and eight active fibers. Electrical mapping, route lengths, cost and power are assumptions here. Optical-budget and polarity qualification remain separate.

Installed capital is 32×switch price +2,048×module price +1,024×channel price=about $2.7M. Spares add 2×switch price +103×module price +21×channel-kit price=about $0.1M. Summing unrounded terms gives about $2.8M, leaving $0.2M; money rounds to $0.1M. Exclusions are servers/NIC electronics, racks, distribution, storage, OOB, taxes and recurring licenses/support. Power is 32×600 W +2,048×12 W=about 44 kW, including host modules. Module heat belongs at its installed location and is counted once.

Two spines per rail replace three: 16/16 uplink bundles save eight switches cluster-wide versus 11/11/10, leaving sixteen rather than twenty-one links after the largest spine loss. The failed-cut test below selects two spines under this deadline. A sole spine leaves no cross-SU path after failure, so no finite blocking ratio makes it eligible.

Stage SU A with sixteen spines, eight leaves and 256 host +256 fabric channels; SU B adds eight leaves and 512 channels without recabling. A spare repairs a leaf after its 32 ranks disconnect. Growth retains eight rails and equal capacity; beyond 1,024 GPUs, new spines require uplink redistribution, reserved panel positions and a maintenance drain.

Staged spine additions on the same 400G radix
Cluster GPUsLeaves / railSpines / railLinks per leaf to each spineTotal leaves / spines
512221616 /16
1,024421632 /16
2,04884864 /32
4,0961684128 /64
8,19232162256 /128
16,38464321512 /256
Every leaf still has 32 downlinks and 32 uplinks. A spine uses 32 ports at 512 GPUs, then 64 at each listed full stage. The final stage has 16,384 host and 16,384 fabric channels, hence 65,536 modules. Rank groups and failure deadlines must be requalified at each stage.

Planning the spine: build for the cluster you intend to finish with

Fabrics are bought once and grown many times, and the two motions are asymmetric: growth lives in the leaves; regret lives in the spine. Adding a scalable unit is additive — new racks, new leaf switches, new cables landing on spine ports that were waiting for them. Widening a spine after the fact is not: every leaf's uplinks must be re-spread across the new spine set to preserve the symmetry the collectives depend on, which means touching cabling fabric-wide on a live cluster. NVIDIA's reference architecture says the quiet part: even when a scalable unit deploys partially, the fabric should be designed and cabled for the full SU, with the unused portion left dark (NVIDIA B300 RA, 2025). The professional pattern is to fix the target scale first, size — and at minimum cable — the spine for it, and let utilization, not architecture, grow with the fleet.

Two sizing levers fall out. Link speed versus spine width: a slower link can require more ports and different spines to preserve the same aggregate, but the result depends on the actual port map. The staged ledger shows which ports remain dark and when existing uplinks must move; no general “half-speed doubles growth” rule replaces that map. Fewer, faster spines can cost less initially while making each lost spine a larger fraction of the surviving uplinks. The unsplit-flow constraint: an unsplit 400G flow cannot sustain its rate on a 200G fabric path. Packet spray can use several links only when the endpoint and transport support it; otherwise provision the path for the flow’s required rate. → Chapter 8.4.

Deliberate overprovisioning — under-subscription — is a named tool, not an accident. Meta's production RoCE fabric, hit by ECMP flow collisions, doubled rail-switch uplink bandwidth to run 1:2 under-subscribed; once better load balancing landed, it trimmed the ratio to 1:1.125 rather than all the way to 1:1 — explicitly keeping buffer for up to two link failures (Meta, SIGCOMM 2024). That is the honest arithmetic of a fabric that must hold its blocking factor through real-world link loss: a few percent of extra uplink capacity, bought once, doubles as failure margin and ECMP smoothing. Rack granularity is the other quiet sizing input: rack-scale systems land NICs in multiples the switch radix may not divide. A GB200 NVL72 rack presents 18 NICs per rail (four GPUs per compute tray, 18 trays, one rail per tray position), so a 64-port leaf running 18 down and 18 up strands 28 ports — the mismatch that made 144-port (8×18) switch radixes the clean fit for NVL72-era fabrics (SemiAnalysis, 2024). Count ports against the rack unit and the rail structure before trusting any aggregate-bandwidth arithmetic; the BOM error hides in the remainders.

Each rank sends W=7.875 GiB. Thirty-two ranks send in each direction, crossing 32×8,455,716,864=270,582,939,648 bytes per direction per rail. Healthy cut service is 32×400/8=1,600 GB/s raw, or 1,440 GB/s useful; transfer time is 270,582,939,648/(1,440×10⁹) seconds. The step is 840 ms at 10 ms precision: pass. Losing Y leaves sixteen links: 800 GB/s raw and 720 GB/s useful. Substitution in 800+max(0, transfer ms−150) gives 1,030 ms at 10 ms precision: pass. Losing all Y-bank spines gives that result on every rail; shared leaves, NICs and power remain common dependencies.

The cut allowance is (1,100−800+150)=450 ms, requiring 270,582,939,648/0.450=about 600 GB/s useful per direction per rail. Division by 45 GB/s per surviving link rounds upward to fourteen links minimum; thirteen give about 1,110 ms: fail. Allowed blocking is 450/(270,582,939,648/(1,440×10⁹)×1,000)≈2.4:1; integer links use the unrounded inequality. The deadline flip occurs at 800+max(0,270,582,939,648/(720×10⁹)×1,000−150) ms, about 1,030 ms. Below it, the two-spine option needs more surviving capacity, fewer bytes or more overlap.

Admission depends on measured service and tail matching the assumption. The first interrupted step remains HOLD until detection, rerouting, retries and repeated compute are tested. A leaf loss disconnects ranks and invokes repair/relocation and restart. Chapter 8.6 owns control/reduction, Chapter 9.4 restart cost and Chapter 13.7 installed acceptance.

The oversubscription fork: start from the published tier, then run three tests

Oversubscription is the ratio of downlink to uplink bandwidth at a tier. A 1:1 tier provisions as much bandwidth toward the spine as it presents to the GPUs — the non-blocking case — while a 3:1 tier provisions one-third as much and blocks the moment more than a third of its GPUs try to leave the tier at once. It is a direct lever on the back-end bill: changing the spine from 1:1 to 2:1 removes capacity and hardware where the port map permits it; this chapter’s illustrative ledger prices those switches and optics, so the saving follows the equipment count. The question is never whether you would like to save that money. It is whether the tier you are blocking carries traffic that will notice.

Published production fabrics show what their workloads could afford to block; their ratios apply when locality and failure budgets match. The cited synchronous-training designs provision non-blocking rail tiers. NVIDIA's SuperPOD references and Meta's production RoCE training fabric both keep the tier that carries the data-parallel all-reduce at 1:1 or better, with Meta at the 1:1.125 under-subscription recorded earlier in this chapter. The same Meta designs run 7:1 at the top layer, where the only traffic is cross-zone and rare (Meta, SIGCOMM 2024). That pair of numbers inside one fabric is the lesson: the ratio is tiered, and two tiers can differ by an order of magnitude because the traffic on them does. Inference designs oversubscribe the upper tier at 2:1–3:1 in the published Juniper and neocloud references, where their measured request placement, concurrency and tail allowances justify that ratio; KV movement and expert routing can instead synchronize a cross-leaf burst.

Three tests move a tier off its starting point, and all three are measurable before you order optics. Locality: what fraction of the dominant collective's bytes leave the tier? A rail-aligned all-reduce stays on the leaf only while the DP group fits one rail switch; beyond that it climbs the pod spine, which is why the published training designs hold the pod spine at 1:1 and block only the top tier that carries cross-pod traffic. By contrast an expert-parallel all-to-all or a disaggregated prefill-to-decode KV transfer crosses rails on every layer or every request, and the tier it crosses must stay near 1:1 (→ Chapter 8.1 for the traffic characterization). Exposure: how much of step time or request latency is communication the compute cannot hide? Where communication is overlapped, a blocked tier costs little until it saturates; where it is exposed, every point of bandwidth lost is a point of step time or a tail-latency miss. Failure headroom: the ratio you buy is not the ratio you run — every dead link or optic on the tier raises the effective ratio for the leaves behind it, which is why Meta's 1:1.125 is a failure budget rather than a rounding error, and why the fabric's optics failure rate (Chapter 8.9) belongs in the sizing model. Run the three tests against the target job's trace or a representative proxy, set each tier, then validate the ratio on the built fabric before the workload arrives.

Oversubscription decision matrix by workload
WorkloadDominant scale-out trafficPublished starting ratioWhat pushes a tier toward 1:1What lets a tier block
Pre-training (synchronous)Data-parallel all-reduce, rail-aligned; pipeline point-to-point1:1 (or under-subscribed) at the rail tier and the pod spine it feeds; block only the top tier carrying cross-pod trafficExposed communication at large DP degree; cross-rail placement; link failures eating the headroomRail-local collectives with overlapped communication; a top tier that carries only cross-pod traffic
Post-training / RLTrainer collectives plus rollout-to-trainer weight sync and sample flowsTrainer tier as pre-training; rollout tier 2:1–3:1Frequent weight sync on a tight staleness budgetRollout pools that tolerate staleness and stream samples, not gradients
Online inferenceRequest routing, KV transfer, expert all-to-all2:1–3:1 at the upper tier in the published designsWide expert parallelism spilling past the scale-up domain; prefill/decode disaggregation moving KV across racksReplicas that stay leaf-local with uncorrelated request peaks
Batch inferenceThroughput-bound request and KV flows with queue tolerance2:1–3:1 and higher where the completion target allowsLittle — only a completion deadline the queue cannot absorbAny blocking the completion objective absorbs; latency is not the SLO
Starting ratios are the published SemiAnalysis, Juniper and Meta designs (2024–2026); the last two columns are the locality, exposure and failure-headroom tests that move a tier off them.

Mapping parallelism onto the topology

A non-blocking fabric is necessary but not sufficient — you also have to place the parallelism on it correctly, or you starve the collective the fabric was built to serve. Match each parallelism dimension's bandwidth appetite to the tier that can feed it, working outward from the fattest pipe.

Keep TP and EP exchanges inside scale-up when their exposed time would exceed the next tier’s budget. These are the most bandwidth-hungry dimensions — TP shards a single layer across GPUs and exchanges activations every layer; EP routes tokens to experts with all-to-all on every MoE layer. The named NVLink 5 and 400G comparison in Chapter 8.1 illustrates the interface gap; measured exposed time determines whether this job needs it. Fit TP and EP inside the scale-up domain when spilling their exchanges from NVLink onto NICs puts the delay on the critical path. A deliberate cross-domain implementation can still pass when capacity requires it and measured message size and overlap leave enough time; an accidental spill pays in lower throughput. This is the deepest reason scale-up domain size is a workload decision — a 72-GPU NVL72 domain admits far wider EP (and thus larger MoE models served efficiently) than an 8-GPU HGX domain. → scale-up fabric in Chapter 8.2; wide-EP economics in Chapter 10.11.

Pipeline parallelism (PP) and data parallelism (DP) ride the scale-out fabric. PP passes activations between pipeline stages — point-to-point, latency-sensitive but not bandwidth-crushing — and DP runs the gradient all-reduce across replicas. Both can use scale-out when the schedule fits: PP overlaps communication with computation only as its microbatches allow, while DP’s bandwidth-heavy all-reduce can follow rail affinity and still cross several leaves and spines. Place DP groups along rails so the all-reduce stays on-rail, and the spine carries both cross-rail paths that need it and same-rail transfers whose endpoints attach to different leaves.

Get the mapping wrong and the damage is invisible in a topology diagram and brutal in production: a perfectly non-blocking fabric, fully provisioned, running a job whose TP dimension landed across the scale-out path — every layer's activation exchange crawling at NIC speed, MFU on the floor, and a network team insisting the fabric is healthy, because it is. Topology-aware scheduling (Chapter 10.2) exists to prevent exactly this; the fabric and the scheduler are co-designed, not independent.

Deep dive: the routed underlay — how the fabric actually forwards

Above the transport sits a routing design, and for Ethernet back-ends one production pattern dominates: a flat routed L3 fabric running BGP (InfiniBand fabrics route via the subnet manager instead). Meta's account of its RoCE clusters needs one line — default routes are provided by BGP — and vendor validated designs converge on the same shape: eBGP (or iBGP with the spines as route reflectors) on every leaf–spine link, ECMP doing the spreading, and no overlay at all on a single-tenant back-end (Meta, SIGCOMM 2024; Juniper validated design; NVIDIA Cumulus reference design). The refinement that makes this tractable across thousands of fabric links is unnumbered BGP: sessions ride the IPv6 link-local addresses interfaces assign themselves, with IPv4 routes advertised over IPv6 next-hops (RFC 8950), so numbered transit subnets can be avoided and peers discovered per interface; endpoint prefixes, router identities, peer policy and failure detection still need a plan, because not every failed forwarding path drops the physical link or BGP session. Multi-tenant fabrics layer VXLAN/EVPN on this underlay where isolation demands it (→ Chapter 8.4).

The frontier-scale counterpoint is instructive: the MRC/SRv6 design deletes dynamic routing from the data plane entirely — switches carry static SRv6 forwarding state configured at install time, and hosts choose paths by stamping segment lists onto packets (OpenAI et al., 2026). The fork is an operating-model choice, not a protocol beauty contest. BGP puts failure response in the fabric's control plane: distributed convergence, no host involvement, and the network team owns the pager. SRv6-static puts it in the host stack and scheduler: endpoint path discovery, impairment detection and per-packet selection around a failed path, at the price of owning path selection in software forever. Decide who you want holding failure response before you pick the underlay.

Routing and admission ownership

For a BGP underlay, start from RFC 7938, August 2016: define leaf/spine roles, AS assignment, endpoint-prefix ownership, allowed advertisements and ECMP policy. RFC 8950, November 2020, carries IPv4 NLRI with IPv6 next hops; it does not advertise IPv4 next hops over an IPv6 session. Pin the NOS’s link-local discovery, multiprotocol capability and next-hop resolution behavior. Keep OOB and tenant management routes outside the back-end export policy. A reachability table without MTU, ECMP and prefix limits is an incomplete forwarding contract.

Who detects the fault and who restores the job?
FaultDetection / forwarding responseOwner / evidence
Local carrier lossPort-down removes the adjacent path; routing updates the usable next-hop setNetwork owns link and FIB recovery; transport owns retry and completion
Remote failed downlinkRemote leaf must withdraw or mark the affected endpoint reachability; a local healthy uplink proves nothingNetwork owns prefix propagation; endpoint/job owner handles an unreachable rank
Carrier-up blackholeBFD tests its configured forwarding path; workload-equivalent probes cover failures BFD does not traverseNetwork owns session/FIB evidence; transport owns loss/error; application owns restart
Host-selected MRC/SRv6 path failureController/endpoint must identify stale or bad paths and update selection; static switch state does not do thisHost path-controller owner plus network inventory owner; prove path choice and job progress
RFC 5880 (June 2010) defines BFD; session health alone does not certify every queue, packet size or endpoint path.

A scheduled replacement adds an admission owner. Meta’s DSF account, 2025-10-20, describes ingress VOQs, egress grants and cell distribution/reassembly. For the same failed-state DP cut, grants must supply the required useful rate in each direction within receiver drain/resource limits. Cap admitted bytes at the actual VOQ allocation and outstanding-credit limit; hold admission at that bound. Withhold a grant and fail a fabric element: require bounded queues, unrelated-destination progress, and correct completion or explicit failure. Reject a configuration whose grants miss the workload deadline. The fabric-system owner owns grants/recovery; the endpoint owner owns completion/retry. Chapter 8.6 derives congestion and reduction; this ledger closes capacity and ownership.

1:1 vs 2:1–3:1
published scale-out examples (1:1, 2:1–3:1, reported 7:1) — workload-specific observations, not defaults; derive from measured traffic, topology and SLO
Scope & caveats

named reference designs and deployments with different traffic, topology, placement and service objectives

Derive oversubscription from the measured traffic matrix, collective/request mix, topology, failure headroom and SLO; validate it on the target fabric.

576 GPUs / SU
NVIDIA GB200 SuperPOD reference SU; separate from the 512-GPU teaching ledger
144 × 800G
Quantum-X800 radix (72 OSFP); 2-tier fat-tree to 10,368 NICs, tens of thousands of GPUs in 3 tiers
102.4 Tbps
Tomahawk 6 chip capacity; family production volume announced 2026-03-12
Scope & caveats

102.4 Tb/s chip family; 512×200G or 1,024×100G electrical options. System ports and optical variants require separate support and delivery evidence.

~18×
NVLink5 scale-up (1.8 TB/s/GPU bidirectional, ~900 GB/s/dir; 130 TB/s NVL72 rack) over ~400G scale-out NIC (~50 GB/s/dir) — keep TP/EP inside scale-up
≤3 hops
dragonfly worst-case diameter; Google OCS cuts a 1,024-chip torus worst case from 16 to ~7 hops
8.4% (10.0% incl. NIC)
network share of Llama 3 interruptions — switch/cable 8.4%; 10.0% including NIC
Scope & caveats

Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.

131,072derived
GPUs in a two-tier fabric of 512×100G switches with each 800G NIC split into eight 100G planes (MRC consortium design; four-plane 4×200G and eight-plane 8×100G systems deployed)
Scope & caveats

Eight-plane design maximum in the May 2026 MRC/SRv6 paper, which reports deployed four-plane (4 × 200G) and eight-plane (8 × 100G) MRC systems.

131,072 = 512 T0 switches × 256 NICs in each of the 8 planes, with every GPU holding one 100G port per plane (design maximum); the consortium paper reports deployed four-plane (4 × 200G) and eight-plane (8 × 100G) systems.

576 x 800G
largest single-chassis AI spine (460 Tbps modular): 576×800G / 1,152×400G — below this, the whole back-end can be a single tier
Scope & caveats

Vendor datasheet figures; the 160k+ two-tier number is a vendor scaling rating, not a documented deployment.

18K x 800G
Meta DSF (VOQ, credit-scheduled, cell-sprayed) reported GPU scale; DES-class products list >27k×800G ports per logical switch
Scope & caveats

DES-class >27k figure is a vendor datasheet maximum (Arista 7700R4), not a deployment report.

1:2 → 1:1.125
Meta rail-switch overprovisioning: 1:2 under-subscribed against flow collisions, later 1:1.125 — keeping buffer for two link failures
38-77%derived
paper-modeled rail-only network-cost reduction vs full-bisection fat-tree
Scope & caveats

Paper-modeled savings, not a production report. All-to-all forwarding overhead differs by paper version (v4: 4.1-5.6%; v5: 8.2-11.2% completion-time overhead) — cite the version. Assumes the collective profile stays rail-aligned.

The cost / reach / blocking tradeoff, made explicit

Every topology decision resolves to a point in a three-axis space — cost, reach, and blocking — and the three are coupled, so you cannot optimize one without paying in the others. Blocking (the inverse of bisection bandwidth you provision) is set by the oversubscription ratio: non-blocking is the most expensive; whether a tier needs it is what the locality, exposure and failure-headroom tests decide. Reach is set by physics: at 800G and 1.6T the copper cliff arrives within a couple of meters, so a floor plan that spreads racks out converts copper links into optical links and multiplies the transceiver count — meaning the building layout is a fabric-cost variable. Cost is the resultant: it rises with switch tiers (more radix to cross), with non-blocking provisioning (more spine), and with optical reach (more transceivers).

The practical optimization is to spend copper where you can and optics where you must, and to push as much traffic as possible into the cheapest tier. Rail-optimized topologies can keep useful rank locality and avoid expensive optical paths only when the port map actually removes connectivity or capacity; a multi-leaf rail still pays for its leaf–spine–leaf traffic. Compact, dense floor plans keep more links inside the copper reach budget. And the SU boundary is drawn, in part, to keep intra-SU links on copper and reserve optics for SU-to-SU. The fabric that looks cheapest on a per-port basis is rarely cheapest in the building; the all-in number is dominated by how many of your links crossed the copper cliff into optics.

Deep dive: why rail-optimized beat flat Clos for synchronous training

A flat non-blocking Clos and a rail-optimized fat-tree can both deliver full bisection bandwidth, so why did rail-optimized become the 2026 default for NVIDIA-class training? Traffic locality. In a synchronous training job the dominant collective — the data-parallel all-reduce — is rail-aligned: GPU i in every node communicates predominantly with GPU i in every other node. A rail-optimized topology wires exactly that pattern onto one rail: GPU i in every node hangs off the switches serving rail i, so the all-reduce that defines the step stays inside a single rail's switch group instead of spreading across the whole fabric. A rail is an endpoint-affinity group; a plane is a separately routed path set. Neither term means one switch. In NVIDIA's GB200 scalable unit each rail is a spine-leaf group of eight leaves — one per compute rack — plus six spines, so the collective stays on a leaf only while it fits one rack, and crosses the spine the moment it spans racks. Size the tier from the rank-to-NIC-to-leaf map, not from the rail label. What rail alignment removes is the cross-rail traffic a flat Clos has to provision for everywhere, so a smaller spine can buy fewer switches and optical links only after the same-rail inter-leaf cut and failure budget still close.

A flat Clos can also host locality-aware placement; a rail-restricted design saves switches or optical links only where that restriction removes connectivity the same traffic and surviving-capacity requirement do not need. Count those deleted paths rather than assuming all-reduce fills every spine link. The trade is generality for efficiency: rail-optimized is optimal for the rail-aligned collective and worse for arbitrary all-to-all — which is why EP-heavy MoE inference, with its all-to-all token routing, leans harder on the scale-up domain and complicates the rail assumption. The general point: the cheapest correct fabric is the one whose topology encodes the dominant collective's communication pattern. → collective and traffic characterization in Chapter 8.1.

Anti-patterns

The recurring fabric-sizing mistakes all share a root cause: sizing from a component spec or a topology preference instead of from the workload's coupling and parallelism plan. Five are worth naming:

  • Buying bisection the traffic cannot use. A non-blocking top tier above pods whose all-reduce never leaves them, or above leaf-local inference replicas, spends money on links whose capacity the tested workload does not use. Run the locality test before you buy the spine.
  • Oversubscribing without a step-time budget. A 2:1 rail tier under a synchronous all-reduce with exposed communication is a step-time tax on every step of every job, paid for the life of the fabric to save a fraction of the optics once. Block the tier whose local traffic leaves spare capacity first; block a tier the dominant collective crosses only when its exposed bytes still meet the deadline with the required failure margin.
  • Parallelism spilling out of the scale-up domain. Placing TP or wide EP across the ~400G scale-out path (~50 GB/s per direction) instead of the ~1.8 TB/s NVLink domain (~900 GB/s per direction) gives the most bandwidth-hungry collective ~18x less bandwidth per direction on a fabric that reports healthy. Overlap and message size mean the step-time penalty is smaller than 18x, and it is still the largest placement error available to you. Keep TP inside the scale-up domain. Expert parallelism is the deliberate exception — hierarchical implementations combine intra-node NVLink with inter-node RDMA precisely because the expert set outgrows one domain — so the error is spilling out by accident and never measuring the step time.
  • Growing the spine reactively. Sizing the spine for this quarter's GPU count and widening it per expansion breaks leaf–spine symmetry and forces fabric-wide recabling on a live cluster. Fix the target scale; build — or at least cable — the spine for it on day one, and grow in leaves and SUs. When the target outruns one fabric's radix, scale with bigger-radix spines or added planes before conceding a third tier.
  • Sizing in GPUs instead of SUs. Ad-hoc GPU counts produce ad-hoc cabling, irregular fault domains, and a BOM that does not replicate. Draw the SU boundary first; size, cable, and procure in SUs.
Traffic characterization and the collectives this chapter sizes for are in Chapter 8.1; the scale-up domain that TP/EP must fit inside is Chapter 8.2; the switch, NIC, and DPU silicon whose radix sets the topology is Chapter 8.3; the protocol and transport layered on this topology is Chapter 8.4. Congestion control and load balancing that keep an oversubscribed fabric honest are Chapter 8.6; the management/OOB fabric and timing are Chapter 8.7; scale-across to multi-campus is Chapter 8.8; the physical-layer reach budget and optic taxonomy that cap the topology are Chapter 8.9, and the fiber plant and structured cabling are Chapter 8.10. The oversubscription fork inherits its logic from the archetype cascade in Chapter 1.1; topology-aware placement of the parallelism plan is scheduled in Chapter 10.2 and exploited for wide-EP inference in Chapter 10.11; fabric commissioning and bisection-bandwidth validation are an acceptance gate in Chapter 13.7.

Choose the least costly counted topology whose actual paths meet the healthy and surviving-state workload budget, then reserve the ports and maintenance plan for the next stage. Reduce spines only while the failed-state cut still passes; a spare on a shelf does not replace a missing live path. Choosing the smaller bill without that test turns a one-time saving into a delay paid by every affected job.

Cite this chapter
Fehn, J. (2026). Scale-Out Topology, Sizing & Oversubscription (Chapter 8.5). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-5-scale-out-topology-sizing-and-oversubscription (accessed 2026-09-29).
@misc{aidc-8-5,
  author       = {Fehn, Jacob},
  title        = {Scale-Out Topology, Sizing & Oversubscription (Chapter 8.5)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-5-scale-out-topology-sizing-and-oversubscription},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit