The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 8.2

In this chapter · 8 sections
Term help

Scale-Up Fabric (Intra-Node / Intra-Rack)

The scale-up domain — accelerators talking at memory speed over a fabric an order of magnitude faster than the back-end — sets your tensor- and expert-parallel ceilings, your MoE economics, and your blast radius.

GOODPUTDENSITY-RAMPPOWER-BOUND

What you'll decide here

  1. How large a supported scale-up domain you actually need — the dated eight-GPU HGX or 72-GPU GB200 configuration, or a future 576-GPU layout — which is a decision about resident state, your TP/EP ceiling and MoE dispatch/combine time. Count GPUs, packages and dies separately before treating rack count as capacity.
  2. Whether you buy the domain as an integrated NVLink/NVSwitch stack, license NVLink Fusion IP, or assemble a qualified UALink/Ethernet scale-up implementation — the lock-in-versus-commoditization fork turns on which endpoint, switch and software interfaces you can actually replace, and therefore who holds supplier power.
  3. Where copper stops and optics start inside the domain — the reach wall is the actual package, board, connector and cable budget, while choosing CPO after that crossing determines which optical failures cost a module, board or switch replacement.
  4. The blast radius you are willing to own: a bigger NVLink domain can lift MFU and widen expert placement, but a switch-tray or link fault can expose many more GPUs to a drain and restart. Distinguish a degraded surviving path from a lost domain in the selected platform’s fault procedure before sizing jobs and spares.
  5. Whether the workload needs peer reads/writes, atomics and ordering within a scale-up domain, coherent access through a supported CXL relationship, or message-passing collectives — misclassifying those operations strands bandwidth on the wrong network. Remote addressability, cache coherence and a collective API are separate contract lines.
Illustrative — stated assumptions. Candidate A supports the required peer and message operations; candidate B lacks a required operation. These are qualification states, not supplier verdicts. Apply the same operation, ordering, topology, procurement and repair axes to the NVIDIA, UALink and Google implementations in the chapter’s matrix. A standard’s endpoint ceiling is not a purchasable service slice, and an ordered peer operation is not evidence of universal coherence. Select a larger domain only when its supported operations remove a demonstrated crossing cost and its repair outage is acceptable.

Chapter 8.1 drew the three-network model — scale-up, scale-out, scale-across — and argued that the boundaries between them move with each silicon generation. This chapter lives inside the innermost ring. The scale-up fabric is the set of accelerators that communicate at memory speed: a switched or topology-specific peer interconnect running roughly eighteen times faster per device, like-for-like per direction, than the back-end NIC, across which a tensor- or expert-parallel shard can be split without the collective collapsing. By 2026 this is no longer an intra-server concern — the domain has climbed out of the eight-GPU server into named designs such as the 72-GPU GB200 NVL72 rack; NVIDIA’s March 2026 Rubin Ultra forecast reaches 576 GPUs across multiple racks, with a distinct topology and software contract.

How big to make it is a decision with a cascade behind it, not a number you read off a datasheet. Too small, and you cap the tensor-parallel group that fits locally or force expert exchanges onto the slower back-end fabric; MoE inference throughput falls when those exposed exchanges miss the request deadline. Too large, and you pay for unused capacity and a broader drain unit that can expose hundreds of GPUs. Routes beyond their electrical budget add an optical bill; choosing CPO then adds its board-or-switch repair boundary and spares cost. Buy it from one vendor and you inherit their roadmap and their margin; assemble it from a standard and you inherit an integration burden and a maturity gap.

What 'scale-up' actually means

The defining property is per-device bandwidth and semantics, not topology or distance. The dated NVIDIA NVLink 5/6 records describe 1.8–3.6 TB/s per accelerator (in NVIDIA's bidirectional-aggregate convention — 900 GB/s to 1.8 TB/s each way, or 14.4–28.8 Tb/s if you would rather think in network units), versus ~50–100 GB/s per direction on a 400–800G scale-out NIC — a ~18x like-for-like gap at both ends of that range, and the entire reason the two networks exist separately. → the bytes-versus-bits and bidirectional-versus-per-direction conversions that decide every comparison in this Part are worked in Chapter 8.1. Inside the scale-up domain, accelerators address each other's HBM with load/store or one-sided semantics and run collectives (all-reduce, all-gather, all-to-all) at a latency and bandwidth the back-end fabric cannot touch. The rule of thumb that falls out of this: fit your tensor parallelism and your expert parallelism inside the scale-up domain; let data and pipeline parallelism span the scale-out fabric. Cross that boundary with the wrong collective and you have moved a memory-bandwidth-bound operation onto a network roughly eighteen times slower.

Domain size is therefore a workload variable, not a ladder of interchangeable maxima. An eight-GPU HGX group and a 72-GPU GB200 NVL72 rack put different amounts of resident model, expert state and workspace inside one local fabric. Count the devices the scheduler can actually allocate together before deciding which TP/EP exchanges must cross the back-end. Then check the operations available across that population: peer reads and writes, the required atomic widths, memory ordering and the runtime’s collective path. Connectivity does not create a coherent shared HBM cache. NVLink-C2C CPU–GPU coherence, CUDA peer access, and MNNVL/IMEX access domains solve different problems. A fast message-passing ICI system belongs in this comparison even though the application does not acquire a general load/store address space. Record what publishes a remote write and who revokes access after process or device failure; without those rules, a larger addressable population is not a usable domain.

NVIDIA's NVLink is the incumbent and the reference against which every challenger is measured, so it anchors the discussion. The architecture is two parts: NVLink, the SerDes-based per-GPU link, and NVSwitch, the crossbar ASIC that turns point-to-point links into an all-to-all switched fabric. Generation by generation, the per-GPU number is the headline: NVLink 4 (Hopper) delivered 900 GB/s, NVLink 5 (Blackwell) doubled it to 1.8 TB/s, and NVLink 6 (Rubin) doubles it again to 3.6 TB/s — over 14x the bandwidth of PCIe Gen6 (NVIDIA, 2026). NVSwitch is what makes a domain out of those links: in GB200 NVL72, nine NVSwitch trays wire 72 Blackwell GPUs into a single non-blocking domain delivering ~130 TB/s of aggregate NVLink bandwidth and ~13.4 TB of aggregate HBM in the GB200 NVL72 profile checked September 2026, accessed and synchronized through the supported programming model rather than one coherent cache.

The NVL domain is the resource the placement and health services must describe. A GB200 NVL72 places 72 GPUs in one rack; keep that record attached to GB200; do not use package or die counts from a successor to enlarge it. The current NVLink Fusion platform page describes an NVLink 6 domain separately from its future domain-size roadmap. Multi-Node NVLink (MNNVL) and Internode Memory Exchange (IMEX) establish supported cross-node access and permissions; they do not turn every connected cache into a coherent copy. The domain is a first-class allocatable resource: a scheduler must intersect that access domain with available GPUs, qualified routes and the job’s recovery unit. Allocating across an unqualified boundary turns a placement convenience into a repeatable communication or protection failure. → Chapter 8.1 for the collective/parallelism mapping.

NVLink-SHARP is an in-domain capability to qualify alongside peer operations. Record the GPU and switch generation, collective library, operation, data type, message range, concurrent-group limit and fallback path. A switch can expose the hardware while a particular call still runs on endpoints. The byte-count comparison and reduction acceptance method have one home in Chapter 8.6; here, select whether the supported domain provides the capability the workload needs.

Buying an NVLink domain ties endpoints, switches and collective software to NVIDIA’s supported interface; that agreement is a strong supplier dependency even when alternative local accelerator fabrics exist. The 2025–26 inflection is that the rest of the industry has converged on open scale-up fabrics, and they have split into a three-way contest. The fork is strategic before it is technical: a supported scale-up domain binds accelerator, switch and software choices and is a major lock-in surface in the stack, because the accelerator, the switch, and the collective library co-design around it. Breaking that open is the explicit goal of the challengers.

NVLink / NVLink Fusion is the integrated or licensed-IP path: Fusion’s 2025-05-18 introduction enables custom endpoint integration, not arbitrary mixed-accelerator interoperation. UALink is an open specification whose 200G 1.0 ratification announcement, 2025-04-08, describes reads, writes and atomics at 200 Gb/s per lane for domains of up to 1,024 accelerators; those specification limits do not prove a purchased system’s population or latency. Its 2026-04-07 family announcement covers Common 2.0, 200G DL/PL 2.0, Manageability 1.0 and Chiplet 1.0; Common adds in-network compute. These are separate specification scopes, not evidence that a purchased combination implements every option. Ethernet scale-up (SUE/ESUN) reuses Ethernet SerDes, switch silicon and interfaces, giving a merchant ASIC potential scale-up and scale-out roles, but the endpoint transport must supply the required memory operations. AMD identifies UALoE in its Helios platform; Ethernet carriage is not native UALink PHY conformance. Broadcom’s 2026-03-12 Tomahawk 6 production-volume announcement establishes silicon status, not the endpoint/software contract for any one rack.

Operations and supported domain: the complete contract
Fabric / implementationSupported population / dateOperationsMemory protectionRepair boundaryMedium / status
NVLink / NVSwitch: GB20072 GPUs; NVIDIA GB200 reference, 2025Peer access and collectives; qualify atomics and orderingCUDA permissions; IMEX when MNNVL is usedSelected platform’s link/switch recovery and drain unitQualified in-rack copper; supported system
NVLink Fusion: NVLink 672 XPUs (platform page, undated)Licensed IP; endpoint generation and software select operationsCustom endpoint plus platform access controlsPartner/OEM procedure; demonstrate degraded routesNamed integration; future 1,152 is separate roadmap
UALink 200G 1.0 / family1,024 accelerator spec limit; 2025-04-08Read/write/atomic; Common 2.0 adds compute, 2026-04-07Versioned permission and address rules to qualifyImplementation’s managed partition and recovery unitPHY conformance and actual system supply are separate
Ethernet scale-up: SUE / UALoEPlatform count, not switch port countRequire endpoint transport’s peer/atomic/ordering setEndpoint authority plus fabric access policyProve path loss, drain and restart for the rackQualified Ethernet channel; no automatic UALink conformance
Google ICI: Ironwood / 8t / 8iDistinct dated chip counts in the records belowRuntime collectives; not general coherent shared HBMService allocation and runtime/device access boundaryService slice, failed links and topology reconfigurationPlatform topology and optical circuits; qualify offered slice
Domain sizes and bandwidths are vendor/consortium records with their publication dates; a maximum population alone does not establish a purchasable, uniformly coherent machine. Medium entries name the qualified electrical route before optics are required. The table carries no bandwidth column; per-device bandwidth is bidirectional aggregate in bytes/s only where the source uses that boundary; multiplying by eight converts bytes to bits (1.8 TB/s = 14.4 Tb/s), not bidirectional aggregate to one-direction service. The conversion belongs to Chapter 8.1.

The table is a bet on supplier power. The vertically-integrated path buys an accelerator, NVLink switch and collective-software stack co-designed and supported as one combination, at the price of single-vendor dependence on the most strategically important boundary in your cluster. The open paths invert that trade: you accept a maturity gap and an integration burden in exchange for multi-vendor switch and accelerator sourcing and a credible threat that keeps incumbent margins honest. Even buying NVLink is a bet — that NVIDIA's roadmap stays far enough ahead, fast enough, to justify the lock-in premium. → the merchant-vs-captive silicon business model is framed in Chapter 8.3.

Adjacent memory-semantic fabrics: CXL and the TPU/ICI alternative

Two adjacent fabrics belong in the same mental model, because their programming and placement contracts differ and neither is a drop-in replacement for a chosen accelerator fabric. CXL (Compute Express Link) is cache-coherent over PCIe physical layers; its sweet spot is memory expansion and pooling — adding or sharing DRAM/HBM capacity across hosts — not the terabyte-per-second all-to-all that training collectives demand. In an AI rack, CXL and the scale-up fabric are complementary: CXL widens the memory pool and disaggregates capacity (increasingly relevant to KV-cache tiering), while NVLink/UALink carries the high-bandwidth collective traffic. Conflating them strands bandwidth on the wrong network — using CXL for an all-reduce, or a scale-up fabric for cold-capacity pooling, both leave performance on the table.

Google’s ICI provides a different scale-up design point: a 3D torus gives a chip six directional neighbors (±X/±Y/±Z), so collective traffic follows topology-dependent hops rather than an all-to-all crossbar. Keep Ironwood’s 2025-04-09 announcement as its own record. Google’s 2026-04-22 architecture account describes an 8t 3D torus and an 8i Boardfly system with different chip populations. Those are platform topology records, not equal service slice sizes. Optical Circuit Switches (OCS) change physical connectivity without becoming packet switches: reconfiguring around faults can avoid losing the same physical neighborhood through repair, but the offered slice shapes and compiler placement determine what a job can exploit. Choose the service on the collective paths and recovery behavior of that slice; a larger advertised population does not promise a shorter critical path. → topology choices recur in Chapter 8.5.

72 XPUs
NVLink 6 / Fusion domain (platform page, undated)
Read the supported endpoint and software scope.
Scope & caveats

Current page’s NVLink 6 domain; future roadmap maximum is a separate claim. Not arbitrary mixed-XPU interoperability.

9,600 chips
TPU 8t torus, announced 2026-04-22
Topology and service slice remain separate.
Scope & caveats

Announced platform topology, not a general coherent memory space or guaranteed service slice availability.

1,152 chips
TPU 8i Boardfly, announced 2026-04-22
Do not rank unlike chip populations as uniform memory.
Scope & caveats

Announced platform topology, not a general coherent memory space or guaranteed service slice availability.

How domain size shapes training and inference

Domain size turns into MFU and tokens-per-dollar here. For training, the scale-up domain sets the ceiling on tensor parallelism and the practical limit on expert parallelism. Tensor parallelism shards a layer’s matmuls across GPUs and exchanges activations according to the selected forward/backward schedule — a latency- and bandwidth-bound all-reduce that stays cheap inside the scale-up fabric when its exposed latency and bandwidth demand would exceed the next tier’s budget. Push TP past the domain boundary onto the back-end network and MFU falls hard, because the per-step collective now runs at scale-out speed. A 72-GPU domain lets a frontier model carry a high TP degree and still leave room for pipeline and data parallelism across the scale-out fabric; an 8-GPU domain holds only TP-8 on the memory-semantic fabric — a wider TP group is implementable, it just pays scale-out latency on every step — so communication outside that group needs a budget for the slower scale-out links it actually traverses.

For inference, the domain is the enabler of wide expert parallelism in Mixture-of-Experts models. Each MoE layer routes tokens to a subset of experts via an all-to-all — the most punishing collective for a slow fabric. NVIDIA's own measurements on NVL72 show wide expert parallelism (EP32 and beyond) substantially outperforming narrow EP8 precisely because the all-to-all stays inside the supported local domain (NVIDIA Developer, 2025). The domain also reshapes prefill/decode disaggregation: a sufficiently large supported domain lets you place a prefill pool and a decode pool in the same NVLink island and stream KV-cache between them at memory speed rather than over the network — the substrate behind GB300/GB200 NVL72 + Dynamo MoE serving. So your maximum profitable EP degree — and with it tokens-per-dollar on large MoE models — depends on resident expert state, dispatch/combine time and recovery cost; a hierarchical implementation can cross scale-up domains when the measured savings justify the slower links. → inference archetype in Chapter 1.3; training archetype in Chapter 1.2.

Resident state is 16×60=960 GiB; sixteen devices reserve 16×64=1,024 GiB. Both placements fit; one eight-device domain alone reserves only 512 GiB and fails. Dispatch bytes are 4,096×2×16×1,024=134,217,728 bytes; combine doubles the transfer to 268,435,456 bytes. In one domain, divide by 300×10⁹ bytes/s: about 0.9 ms, within 2 ms. Split evenly, half crosses the cut: 134,217,728/(45×10⁹) s is about 3 ms before any local-transfer delay, so fail. Select the 16-device domain and accept its larger scheduling and drain unit.

The expert-count flip occurs at eight. With E≤8 and unchanged 60 GiB/expert, E×60≤512 GiB and one eight-device domain fits; the assumed local service still finishes the same token traffic within the allowance. At E=9, 540 GiB exceeds 512 GiB, so the smaller domain fails capacity even before communication. If E exceeds sixteen, the selected local domain also loses its one-expert-per-device placement: acquire a larger supported domain or qualify hierarchical EP with a larger time allowance or faster cut. The NVIDIA Hybrid-EP design, 2026, supplies the mechanism, not these rates. Chapter 8.1 owns the dependency budget; Chapter 8.5 prices the cut.

1.8 → 3.6 TB/s
NVLink per-GPU bandwidth: Gen5 (Blackwell) → Gen6 (Rubin); Gen4 (Hopper) was 900 GB/s
Scope & caveats

Named NVIDIA generation/platform, bidirectional aggregate convention; not delivered collective bandwidth, cache coherence or arbitrary mixed-platform performance.

~130 TB/s
GB200 NVL72 aggregate bidirectional NVLink bandwidth (72 × 1.8 TB/s); one-way injection is half of it, and aggregate HBM is not a coherence guarantee
Scope & caveats

Named NVIDIA generation/platform, bidirectional aggregate convention; not delivered collective bandwidth, cache coherence or arbitrary mixed-platform performance.

up to 1,152 XPUs (future roadmap)forecast
NVLink Fusion future roadmap population; no delivery date
Scope & caveats

Future roadmap population, separate from the page’s NVLink 6 72-XPU domain. No purchaser availability implied.

1,024
UALink 200G 1.0 specification limit, ratification announced 2025-04-08
Scope & caveats

Specification limit of 1,024 accelerators with 200G/lane connectivity; not a qualified purchasable system or a latency/reach guarantee.

512 XPUs
single-hop all-to-all scale-up domain on one Broadcom Tomahawk 6 (102.4 Tbps, SUE/ESUN)
9,216 chips
Ironwood announced superpod chip population, 2025-04-09; separate from TPU 8t/8i
Scope & caveats

9,216 chips in the announced Ironwood platform; not a generic coherent-memory domain or the TPU8t/8i topology.

Partition, degrade, drain and restart

Define three boundaries before allocating a domain: the peers authorized to map memory, the routes healthy enough to carry the job, and the resources that must drain for a repair. On GB200 with MNNVL, NVIDIA’s IMEX/Kubernetes design makes the access domain explicit; it is not a substitute for health-aware placement. A failed link first removes a path. Continue only if the platform supports that degraded state and the remaining paths still connect every placed rank with the required protection and time budget. A switch fault can remove several paths at once; derive the surviving graph from that platform’s topology instead of multiplying a generic “GPUs affected” figure.

When that graph or the supported service procedure fails, stop new placement into the affected unit, quiesce surviving ranks, checkpoint only if the application can still produce a consistent checkpoint, revoke stale peer mappings and drain the required unit. Restart from the last valid checkpoint when quiescence cannot complete. After repair, revalidate link health, access controls and the same collective before returning the unit to the allocator. Count lost useful work and the drain/restart time in Chapter 9.4; Chapter 13.7 proves the installed fault cases. A link count alone does not authorize service during live operation.

Copper vs optical inside the domain — and the CPO transition

The physical gate is the qualified electrical route. Package escape, host-board trace, connector transitions and cable length all consume it. GB200 NVL72’s compact copper spine keeps its in-rack NVLink routes electrical; avoiding optical engines and lasers saves their power and removes their repair exposure, which is why copper is a deliberate strength when the qualified route fits. But no port-rate label supplies a universal distance: retain the exact endpoint generation, cable assembly, route and thermal envelope that passed qualification. When that route fails, compare a shorter layout, active copper and supported optics on delivered service and repair time. Chapter 8.9 owns the channel budget rather than a second reach derivation here.

The wall arrives when the domain's routes outgrow qualified copper. NVIDIA's 2026-03-16 design account forecasts a 576-GPU Vera Rubin Ultra NVL576 domain across eight 72-GPU MGX racks with copper and direct optical connections; that named layout makes rack-to-rack optics part of the design, not a rule for every multi-rack domain. That is the reach argument behind optical scale-up, while co-packaged optics (CPO) is a separate packaging choice. CPO moves optical engines alongside the switch ASIC, shortening electrical paths; its power saving depends on the matched engine, laser and pluggable alternatives, and the saving multiplies with the installed link count. The NVL576 design announcement establishes a future optical topology, not a qualified CPO-NVSwitch shipment date. Broadcom's 2025-10-08 Davisson announcement specifies early-access sampling; NVIDIA separately declared Spectrum-X Ethernet Photonics in production on 2026-05-31, an Ethernet status that does not establish NVLink optical availability. A domain whose routes exceed its copper budget commits the buyer to optics, but the reach decision and the packaging decision have different consequences. A supported span can use faceplate pluggables, a near-package engine or co-packaged optics; moving the engine off the faceplate changes which failures need a module, board or switch replacement and therefore changes repair time and spares. → CPO and the fiber plant are engineered in Chapter 8.10.

Deep dive: why GB200 NVL72 stayed copper and the NVL576 design adds optics

The GB200 NVL72 is a concrete example of keeping a domain inside copper's reach, and understanding it explains why longer routes change the medium decision. In NVIDIA's 2024-10-15 OCP contribution, the rack holds 72 GPUs and nine switch trays connected through copper NVLink cartridges containing over 5,000 cables. The engineering rationale has three parts. Power: keeping those links electrical avoids the optical engines and laser load that an optical alternative would add to the same rack's heat balance. Reliability: avoiding those optical components removes their failure and repair exposure from the synchronous job, while the copper cartridges, connectors and switches retain their own fault paths. Cost: the copper design avoids buying an optical interface at every internal link end; the saving depends on the matched component prices and repair boundary. The NVL72 mechanical design keeps its stated 130 TB/s aggregate bandwidth on a qualified copper spine; a denser layout that makes a failed connector inaccessible exchanges part of that component saving for a longer repair.

The copper envelope must be checked again when a domain spreads across racks. NVIDIA’s March 2026 future NVL576 layout distributes 576 GPUs across eight 72-GPU MGX racks and includes direct optical connections; the longer routes explain the medium choice, while optical-engine placement separately sets power and repair cost. Domain size, medium and rack power are coupled: additional local memory or lower exposed communication must earn the extra infrastructure. Reserve pathways and a service route where that option is valuable. Chapter 8.9 closes signal margin, Chapter 8.10 closes fibers, polarity and replacement boundaries, and Chapter 16.2 owns the dated future rack and platform profiles.

The scale-up roadmap, and where it points

The trajectory toward larger domains is tightly coupled to power, but each platform supplies its own planning boundary. NVIDIA's NVLink platform page assigns 900 GB/s, 1.8 TB/s and 3.6 TB/s per GPU to Hopper, Blackwell and Rubin respectively; that sequence is a set of platform specifications, not a law that every generation doubles useful collective throughput. Domain size follows the same distinction: the 2026-03-16 design account describes Vera Rubin NVL72 separately from future Rubin Ultra NVL576 across eight MGX racks and Kyber's future NVL144/NVL1152 layouts. The medium moves from copper to optics where those named routes require it; the rack boundary alone does not select CPO. The power envelope belongs to the same purchased profile: HPE's GB200 and GB300 rack inputs, its VR200 cabinet basis, and future Kyber power/distribution plans cannot be joined into one interchangeable rack specification. These three curves — bandwidth, domain size and medium — constrain one another, so a roadmap that treats them independently mis-sizes power, cooling and pathways. The multi-rack step also commits the buyer to a larger scheduling and recovery boundary; the future topology becomes a procurement option only when its endpoint, switch, software and repair combination meets the project's qualification date. Space and pathways preserve that option, while an equipment order against an unqualified delivery forecast puts the deployment schedule at risk. The dated subsystem profiles, including future high-voltage distribution and optical-domain timing, stay in Chapter 16.2; the macro power-bound rationale is in Chapter 16.1.

The collective primitives and parallelism mapping that justify fitting TP/EP inside the scale-up domain are in Chapter 8.1. The switch ASICs, NICs, and the merchant-vs-captive silicon business model behind these fabrics are in Chapter 8.3. Scale-out transport and the InfiniBand-vs-Ethernet-vs-Ultra-Ethernet contest pick up where this chapter's boundary ends in Chapter 8.4; topology and oversubscription in Chapter 8.5; the canonical in-network reduction method, including SHARP, and congestion control in Chapter 8.6; multi-campus scale-across in Chapter 8.8. The physical-layer reach taxonomy (DAC/AEC/optics) is Chapter 8.9, and CPO plus the fiber plant is Chapter 8.10. The blast-radius and checkpoint coupling is engineered in Chapter 9.4; the rack-power and 800 VDC substrate this fabric rides on is Chapter 16.2.

Choose the smallest supported domain that fits the placed state and meets the exposed-communication budget with its required operations intact. Pay for a larger domain when it removes a demonstrated deadline miss; accept hierarchical crossings when the local capacity limit forces them and the tested cut can carry them. The alternative is a predictable cost in idle accelerators, constrained placement or a larger restart.

Cite this chapter
Fehn, J. (2026). Scale-Up Fabric (Intra-Node / Intra-Rack) (Chapter 8.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-2-scale-up-fabric-intra-node-intra-rack (accessed 2026-09-29).
@misc{aidc-8-2,
  author       = {Fehn, Jacob},
  title        = {Scale-Up Fabric (Intra-Node / Intra-Rack) (Chapter 8.2)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-2-scale-up-fabric-intra-node-intra-rack},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit