Chapter 8.2
In this chapter · 8 sections
Scale-Up Fabric (Intra-Node / Intra-Rack)
The scale-up domain — accelerators talking at memory speed over a fabric an order of magnitude faster than the back-end — sets your tensor- and expert-parallel ceilings, your MoE economics, and your blast radius.
What you'll decide here
- How large a supported scale-up domain you actually need — the dated eight-GPU HGX or 72-GPU GB200 configuration, or a future 576-GPU layout — which is a decision about resident state, your TP/EP ceiling and MoE dispatch/combine time. Count GPUs, packages and dies separately before treating rack count as capacity.
- Whether you buy the domain as an integrated NVLink/NVSwitch stack, license NVLink Fusion IP, or assemble a qualified UALink/Ethernet scale-up implementation — the lock-in-versus-commoditization fork turns on which endpoint, switch and software interfaces you can actually replace, and therefore who holds supplier power.
- Where copper stops and optics start inside the domain — the reach wall is the actual package, board, connector and cable budget, while choosing CPO after that crossing determines which optical failures cost a module, board or switch replacement.
- The blast radius you are willing to own: a bigger NVLink domain can lift MFU and widen expert placement, but a switch-tray or link fault can expose many more GPUs to a drain and restart. Distinguish a degraded surviving path from a lost domain in the selected platform’s fault procedure before sizing jobs and spares.
- Whether the workload needs peer reads/writes, atomics and ordering within a scale-up domain, coherent access through a supported CXL relationship, or message-passing collectives — misclassifying those operations strands bandwidth on the wrong network. Remote addressability, cache coherence and a collective API are separate contract lines.
Chapter 8.1 drew the three-network model — scale-up, scale-out, scale-across — and argued that the boundaries between them move with each silicon generation. This chapter lives inside the innermost ring. The scale-up fabric is the set of accelerators that communicate at memory speed: a switched or topology-specific peer interconnect running roughly eighteen times faster per device, like-for-like per direction, than the back-end NIC, across which a tensor- or expert-parallel shard can be split without the collective collapsing. By 2026 this is no longer an intra-server concern — the domain has climbed out of the eight-GPU server into named designs such as the 72-GPU GB200 NVL72 rack; NVIDIA’s March 2026 Rubin Ultra forecast reaches 576 GPUs across multiple racks, with a distinct topology and software contract.
How big to make it is a decision with a cascade behind it, not a number you read off a datasheet. Too small, and you cap the tensor-parallel group that fits locally or force expert exchanges onto the slower back-end fabric; MoE inference throughput falls when those exposed exchanges miss the request deadline. Too large, and you pay for unused capacity and a broader drain unit that can expose hundreds of GPUs. Routes beyond their electrical budget add an optical bill; choosing CPO then adds its board-or-switch repair boundary and spares cost. Buy it from one vendor and you inherit their roadmap and their margin; assemble it from a standard and you inherit an integration burden and a maturity gap.
What 'scale-up' actually means
The defining property is per-device bandwidth and semantics, not topology or distance. The dated NVIDIA NVLink 5/6 records describe 1.8–3.6 TB/s per accelerator (in NVIDIA's bidirectional-aggregate convention — 900 GB/s to 1.8 TB/s each way, or 14.4–28.8 Tb/s if you would rather think in network units), versus ~50–100 GB/s per direction on a 400–800G scale-out NIC — a ~18x like-for-like gap at both ends of that range, and the entire reason the two networks exist separately. → the bytes-versus-bits and bidirectional-versus-per-direction conversions that decide every comparison in this Part are worked in Chapter 8.1. Inside the scale-up domain, accelerators address each other's HBM with load/store or one-sided semantics and run collectives (all-reduce, all-gather, all-to-all) at a latency and bandwidth the back-end fabric cannot touch. The rule of thumb that falls out of this: fit your tensor parallelism and your expert parallelism inside the scale-up domain; let data and pipeline parallelism span the scale-out fabric. Cross that boundary with the wrong collective and you have moved a memory-bandwidth-bound operation onto a network roughly eighteen times slower.
Domain size is therefore a workload variable, not a ladder of interchangeable maxima. An eight-GPU HGX group and a 72-GPU GB200 NVL72 rack put different amounts of resident model, expert state and workspace inside one local fabric. Count the devices the scheduler can actually allocate together before deciding which TP/EP exchanges must cross the back-end. Then check the operations available across that population: peer reads and writes, the required atomic widths, memory ordering and the runtime’s collective path. Connectivity does not create a coherent shared HBM cache. NVLink-C2C CPU–GPU coherence, CUDA peer access, and MNNVL/IMEX access domains solve different problems. A fast message-passing ICI system belongs in this comparison even though the application does not acquire a general load/store address space. Record what publishes a remote write and who revokes access after process or device failure; without those rules, a larger addressable population is not a usable domain.
NVLink, NVSwitch, and the NVL domain
NVIDIA's NVLink is the incumbent and the reference against which every challenger is measured, so it anchors the discussion. The architecture is two parts: NVLink, the SerDes-based per-GPU link, and NVSwitch, the crossbar ASIC that turns point-to-point links into an all-to-all switched fabric. Generation by generation, the per-GPU number is the headline: NVLink 4 (Hopper) delivered 900 GB/s, NVLink 5 (Blackwell) doubled it to 1.8 TB/s, and NVLink 6 (Rubin) doubles it again to 3.6 TB/s — over 14x the bandwidth of PCIe Gen6 (NVIDIA, 2026). NVSwitch is what makes a domain out of those links: in GB200 NVL72, nine NVSwitch trays wire 72 Blackwell GPUs into a single non-blocking domain delivering ~130 TB/s of aggregate NVLink bandwidth and ~13.4 TB of aggregate HBM in the GB200 NVL72 profile checked September 2026, accessed and synchronized through the supported programming model rather than one coherent cache.
The NVL domain is the resource the placement and health services must describe. A GB200 NVL72 places 72 GPUs in one rack; keep that record attached to GB200; do not use package or die counts from a successor to enlarge it. The current NVLink Fusion platform page describes an NVLink 6 domain separately from its future domain-size roadmap. Multi-Node NVLink (MNNVL) and Internode Memory Exchange (IMEX) establish supported cross-node access and permissions; they do not turn every connected cache into a coherent copy. The domain is a first-class allocatable resource: a scheduler must intersect that access domain with available GPUs, qualified routes and the job’s recovery unit. Allocating across an unqualified boundary turns a placement convenience into a repeatable communication or protection failure. → Chapter 8.1 for the collective/parallelism mapping.
NVLink-SHARP is an in-domain capability to qualify alongside peer operations. Record the GPU and switch generation, collective library, operation, data type, message range, concurrent-group limit and fallback path. A switch can expose the hardware while a particular call still runs on endpoints. The byte-count comparison and reduction acceptance method have one home in Chapter 8.6; here, select whether the supported domain provides the capability the workload needs.
NVLink, UALink and Ethernet scale-up: operations and supported domain
Buying an NVLink domain ties endpoints, switches and collective software to NVIDIA’s supported interface; that agreement is a strong supplier dependency even when alternative local accelerator fabrics exist. The 2025–26 inflection is that the rest of the industry has converged on open scale-up fabrics, and they have split into a three-way contest. The fork is strategic before it is technical: a supported scale-up domain binds accelerator, switch and software choices and is a major lock-in surface in the stack, because the accelerator, the switch, and the collective library co-design around it. Breaking that open is the explicit goal of the challengers.
NVLink / NVLink Fusion is the integrated or licensed-IP path: Fusion’s 2025-05-18 introduction enables custom endpoint integration, not arbitrary mixed-accelerator interoperation. UALink is an open specification whose 200G 1.0 ratification announcement, 2025-04-08, describes reads, writes and atomics at 200 Gb/s per lane for domains of up to 1,024 accelerators; those specification limits do not prove a purchased system’s population or latency. Its 2026-04-07 family announcement covers Common 2.0, 200G DL/PL 2.0, Manageability 1.0 and Chiplet 1.0; Common adds in-network compute. These are separate specification scopes, not evidence that a purchased combination implements every option. Ethernet scale-up (SUE/ESUN) reuses Ethernet SerDes, switch silicon and interfaces, giving a merchant ASIC potential scale-up and scale-out roles, but the endpoint transport must supply the required memory operations. AMD identifies UALoE in its Helios platform; Ethernet carriage is not native UALink PHY conformance. Broadcom’s 2026-03-12 Tomahawk 6 production-volume announcement establishes silicon status, not the endpoint/software contract for any one rack.
| Fabric / implementation | Supported population / date | Operations | Memory protection | Repair boundary | Medium / status |
|---|---|---|---|---|---|
| NVLink / NVSwitch: GB200 | 72 GPUs; NVIDIA GB200 reference, 2025 | Peer access and collectives; qualify atomics and ordering | CUDA permissions; IMEX when MNNVL is used | Selected platform’s link/switch recovery and drain unit | Qualified in-rack copper; supported system |
| NVLink Fusion: NVLink 6 | 72 XPUs (platform page, undated) | Licensed IP; endpoint generation and software select operations | Custom endpoint plus platform access controls | Partner/OEM procedure; demonstrate degraded routes | Named integration; future 1,152 is separate roadmap |
| UALink 200G 1.0 / family | 1,024 accelerator spec limit; 2025-04-08 | Read/write/atomic; Common 2.0 adds compute, 2026-04-07 | Versioned permission and address rules to qualify | Implementation’s managed partition and recovery unit | PHY conformance and actual system supply are separate |
| Ethernet scale-up: SUE / UALoE | Platform count, not switch port count | Require endpoint transport’s peer/atomic/ordering set | Endpoint authority plus fabric access policy | Prove path loss, drain and restart for the rack | Qualified Ethernet channel; no automatic UALink conformance |
| Google ICI: Ironwood / 8t / 8i | Distinct dated chip counts in the records below | Runtime collectives; not general coherent shared HBM | Service allocation and runtime/device access boundary | Service slice, failed links and topology reconfiguration | Platform topology and optical circuits; qualify offered slice |
The table is a bet on supplier power. The vertically-integrated path buys an accelerator, NVLink switch and collective-software stack co-designed and supported as one combination, at the price of single-vendor dependence on the most strategically important boundary in your cluster. The open paths invert that trade: you accept a maturity gap and an integration burden in exchange for multi-vendor switch and accelerator sourcing and a credible threat that keeps incumbent margins honest. Even buying NVLink is a bet — that NVIDIA's roadmap stays far enough ahead, fast enough, to justify the lock-in premium. → the merchant-vs-captive silicon business model is framed in Chapter 8.3.
Adjacent memory-semantic fabrics: CXL and the TPU/ICI alternative
Two adjacent fabrics belong in the same mental model, because their programming and placement contracts differ and neither is a drop-in replacement for a chosen accelerator fabric. CXL (Compute Express Link) is cache-coherent over PCIe physical layers; its sweet spot is memory expansion and pooling — adding or sharing DRAM/HBM capacity across hosts — not the terabyte-per-second all-to-all that training collectives demand. In an AI rack, CXL and the scale-up fabric are complementary: CXL widens the memory pool and disaggregates capacity (increasingly relevant to KV-cache tiering), while NVLink/UALink carries the high-bandwidth collective traffic. Conflating them strands bandwidth on the wrong network — using CXL for an all-reduce, or a scale-up fabric for cold-capacity pooling, both leave performance on the table.
Google’s ICI provides a different scale-up design point: a 3D torus gives a chip six directional neighbors (±X/±Y/±Z), so collective traffic follows topology-dependent hops rather than an all-to-all crossbar. Keep Ironwood’s 2025-04-09 announcement as its own record. Google’s 2026-04-22 architecture account describes an 8t 3D torus and an 8i Boardfly system with different chip populations. Those are platform topology records, not equal service slice sizes. Optical Circuit Switches (OCS) change physical connectivity without becoming packet switches: reconfiguring around faults can avoid losing the same physical neighborhood through repair, but the offered slice shapes and compiler placement determine what a job can exploit. Choose the service on the collective paths and recovery behavior of that slice; a larger advertised population does not promise a shorter critical path. → topology choices recur in Chapter 8.5.
Scope & caveats
Current page’s NVLink 6 domain; future roadmap maximum is a separate claim. Not arbitrary mixed-XPU interoperability.
Scope & caveats
Announced platform topology, not a general coherent memory space or guaranteed service slice availability.
Scope & caveats
Announced platform topology, not a general coherent memory space or guaranteed service slice availability.
How domain size shapes training and inference
Domain size turns into MFU and tokens-per-dollar here. For training, the scale-up domain sets the ceiling on tensor parallelism and the practical limit on expert parallelism. Tensor parallelism shards a layer’s matmuls across GPUs and exchanges activations according to the selected forward/backward schedule — a latency- and bandwidth-bound all-reduce that stays cheap inside the scale-up fabric when its exposed latency and bandwidth demand would exceed the next tier’s budget. Push TP past the domain boundary onto the back-end network and MFU falls hard, because the per-step collective now runs at scale-out speed. A 72-GPU domain lets a frontier model carry a high TP degree and still leave room for pipeline and data parallelism across the scale-out fabric; an 8-GPU domain holds only TP-8 on the memory-semantic fabric — a wider TP group is implementable, it just pays scale-out latency on every step — so communication outside that group needs a budget for the slower scale-out links it actually traverses.
For inference, the domain is the enabler of wide expert parallelism in Mixture-of-Experts models. Each MoE layer routes tokens to a subset of experts via an all-to-all — the most punishing collective for a slow fabric. NVIDIA's own measurements on NVL72 show wide expert parallelism (EP32 and beyond) substantially outperforming narrow EP8 precisely because the all-to-all stays inside the supported local domain (NVIDIA Developer, 2025). The domain also reshapes prefill/decode disaggregation: a sufficiently large supported domain lets you place a prefill pool and a decode pool in the same NVLink island and stream KV-cache between them at memory speed rather than over the network — the substrate behind GB300/GB200 NVL72 + Dynamo MoE serving. So your maximum profitable EP degree — and with it tokens-per-dollar on large MoE models — depends on resident expert state, dispatch/combine time and recovery cost; a hierarchical implementation can cross scale-up domains when the measured savings justify the slower links. → inference archetype in Chapter 1.3; training archetype in Chapter 1.2.
Resident state is 16×60=960 GiB; sixteen devices reserve 16×64=1,024 GiB. Both placements fit; one eight-device domain alone reserves only 512 GiB and fails. Dispatch bytes are 4,096×2×16×1,024=134,217,728 bytes; combine doubles the transfer to 268,435,456 bytes. In one domain, divide by 300×10⁹ bytes/s: about 0.9 ms, within 2 ms. Split evenly, half crosses the cut: 134,217,728/(45×10⁹) s is about 3 ms before any local-transfer delay, so fail. Select the 16-device domain and accept its larger scheduling and drain unit.
The expert-count flip occurs at eight. With E≤8 and unchanged 60 GiB/expert, E×60≤512 GiB and one eight-device domain fits; the assumed local service still finishes the same token traffic within the allowance. At E=9, 540 GiB exceeds 512 GiB, so the smaller domain fails capacity even before communication. If E exceeds sixteen, the selected local domain also loses its one-expert-per-device placement: acquire a larger supported domain or qualify hierarchical EP with a larger time allowance or faster cut. The NVIDIA Hybrid-EP design, 2026, supplies the mechanism, not these rates. Chapter 8.1 owns the dependency budget; Chapter 8.5 prices the cut.
Scope & caveats
Named NVIDIA generation/platform, bidirectional aggregate convention; not delivered collective bandwidth, cache coherence or arbitrary mixed-platform performance.
Scope & caveats
Named NVIDIA generation/platform, bidirectional aggregate convention; not delivered collective bandwidth, cache coherence or arbitrary mixed-platform performance.
Scope & caveats
Future roadmap population, separate from the page’s NVLink 6 72-XPU domain. No purchaser availability implied.
Scope & caveats
Specification limit of 1,024 accelerators with 200G/lane connectivity; not a qualified purchasable system or a latency/reach guarantee.
Scope & caveats
9,216 chips in the announced Ironwood platform; not a generic coherent-memory domain or the TPU8t/8i topology.
Partition, degrade, drain and restart
Define three boundaries before allocating a domain: the peers authorized to map memory, the routes healthy enough to carry the job, and the resources that must drain for a repair. On GB200 with MNNVL, NVIDIA’s IMEX/Kubernetes design makes the access domain explicit; it is not a substitute for health-aware placement. A failed link first removes a path. Continue only if the platform supports that degraded state and the remaining paths still connect every placed rank with the required protection and time budget. A switch fault can remove several paths at once; derive the surviving graph from that platform’s topology instead of multiplying a generic “GPUs affected” figure.
When that graph or the supported service procedure fails, stop new placement into the affected unit, quiesce surviving ranks, checkpoint only if the application can still produce a consistent checkpoint, revoke stale peer mappings and drain the required unit. Restart from the last valid checkpoint when quiescence cannot complete. After repair, revalidate link health, access controls and the same collective before returning the unit to the allocator. Count lost useful work and the drain/restart time in Chapter 9.4; Chapter 13.7 proves the installed fault cases. A link count alone does not authorize service during live operation.
Copper vs optical inside the domain — and the CPO transition
The physical gate is the qualified electrical route. Package escape, host-board trace, connector transitions and cable length all consume it. GB200 NVL72’s compact copper spine keeps its in-rack NVLink routes electrical; avoiding optical engines and lasers saves their power and removes their repair exposure, which is why copper is a deliberate strength when the qualified route fits. But no port-rate label supplies a universal distance: retain the exact endpoint generation, cable assembly, route and thermal envelope that passed qualification. When that route fails, compare a shorter layout, active copper and supported optics on delivered service and repair time. Chapter 8.9 owns the channel budget rather than a second reach derivation here.
The wall arrives when the domain's routes outgrow qualified copper. NVIDIA's 2026-03-16 design account forecasts a 576-GPU Vera Rubin Ultra NVL576 domain across eight 72-GPU MGX racks with copper and direct optical connections; that named layout makes rack-to-rack optics part of the design, not a rule for every multi-rack domain. That is the reach argument behind optical scale-up, while co-packaged optics (CPO) is a separate packaging choice. CPO moves optical engines alongside the switch ASIC, shortening electrical paths; its power saving depends on the matched engine, laser and pluggable alternatives, and the saving multiplies with the installed link count. The NVL576 design announcement establishes a future optical topology, not a qualified CPO-NVSwitch shipment date. Broadcom's 2025-10-08 Davisson announcement specifies early-access sampling; NVIDIA separately declared Spectrum-X Ethernet Photonics in production on 2026-05-31, an Ethernet status that does not establish NVLink optical availability. A domain whose routes exceed its copper budget commits the buyer to optics, but the reach decision and the packaging decision have different consequences. A supported span can use faceplate pluggables, a near-package engine or co-packaged optics; moving the engine off the faceplate changes which failures need a module, board or switch replacement and therefore changes repair time and spares. → CPO and the fiber plant are engineered in Chapter 8.10.
Deep dive: why GB200 NVL72 stayed copper and the NVL576 design adds optics
The GB200 NVL72 is a concrete example of keeping a domain inside copper's reach, and understanding it explains why longer routes change the medium decision. In NVIDIA's 2024-10-15 OCP contribution, the rack holds 72 GPUs and nine switch trays connected through copper NVLink cartridges containing over 5,000 cables. The engineering rationale has three parts. Power: keeping those links electrical avoids the optical engines and laser load that an optical alternative would add to the same rack's heat balance. Reliability: avoiding those optical components removes their failure and repair exposure from the synchronous job, while the copper cartridges, connectors and switches retain their own fault paths. Cost: the copper design avoids buying an optical interface at every internal link end; the saving depends on the matched component prices and repair boundary. The NVL72 mechanical design keeps its stated 130 TB/s aggregate bandwidth on a qualified copper spine; a denser layout that makes a failed connector inaccessible exchanges part of that component saving for a longer repair.
The copper envelope must be checked again when a domain spreads across racks. NVIDIA’s March 2026 future NVL576 layout distributes 576 GPUs across eight 72-GPU MGX racks and includes direct optical connections; the longer routes explain the medium choice, while optical-engine placement separately sets power and repair cost. Domain size, medium and rack power are coupled: additional local memory or lower exposed communication must earn the extra infrastructure. Reserve pathways and a service route where that option is valuable. Chapter 8.9 closes signal margin, Chapter 8.10 closes fibers, polarity and replacement boundaries, and Chapter 16.2 owns the dated future rack and platform profiles.
The scale-up roadmap, and where it points
The trajectory toward larger domains is tightly coupled to power, but each platform supplies its own planning boundary. NVIDIA's NVLink platform page assigns 900 GB/s, 1.8 TB/s and 3.6 TB/s per GPU to Hopper, Blackwell and Rubin respectively; that sequence is a set of platform specifications, not a law that every generation doubles useful collective throughput. Domain size follows the same distinction: the 2026-03-16 design account describes Vera Rubin NVL72 separately from future Rubin Ultra NVL576 across eight MGX racks and Kyber's future NVL144/NVL1152 layouts. The medium moves from copper to optics where those named routes require it; the rack boundary alone does not select CPO. The power envelope belongs to the same purchased profile: HPE's GB200 and GB300 rack inputs, its VR200 cabinet basis, and future Kyber power/distribution plans cannot be joined into one interchangeable rack specification. These three curves — bandwidth, domain size and medium — constrain one another, so a roadmap that treats them independently mis-sizes power, cooling and pathways. The multi-rack step also commits the buyer to a larger scheduling and recovery boundary; the future topology becomes a procurement option only when its endpoint, switch, software and repair combination meets the project's qualification date. Space and pathways preserve that option, while an equipment order against an unqualified delivery forecast puts the deployment schedule at risk. The dated subsystem profiles, including future high-voltage distribution and optical-domain timing, stay in Chapter 16.2; the macro power-bound rationale is in Chapter 16.1.
Choose the smallest supported domain that fits the placed state and meets the exposed-communication budget with its required operations intact. Pay for a larger domain when it removes a demonstrated deadline miss; accept hierarchical crossings when the local capacity limit forces them and the tested cut can carry them. The alternative is a predictable cost in idle accelerators, constrained placement or a larger restart.
Cite this chapter
Fehn, J. (2026). Scale-Up Fabric (Intra-Node / Intra-Rack) (Chapter 8.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-2-scale-up-fabric-intra-node-intra-rack (accessed 2026-09-29).
@misc{aidc-8-2,
author = {Fehn, Jacob},
title = {Scale-Up Fabric (Intra-Node / Intra-Rack) (Chapter 8.2)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-2-scale-up-fabric-intra-node-intra-rack},
note = {Accessed 2026-09-29}
}