The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuidePart 8

Part 8

Networking, Fabrics & Optics

10 chapters

8.1
Network Fundamentals & AI Traffic Characterization
In a synchronous AI cluster the network sets the pace of every accelerator: the job runs at the speed of its slowest collective and stalls on its longest tail.
8.2
Scale-Up Fabric (Intra-Node / Intra-Rack)
The scale-up domain — accelerators talking at memory speed over a fabric an order of magnitude faster than the back-end — sets your tensor- and expert-parallel ceilings, your MoE economics, and your blast radius.
8.3
Network Silicon: Switch ASICs, NICs & DPUs
The switch ASIC, the NIC, and the DPU set the ceiling on every fabric above them, so their SerDes generation, buffer architecture, and offload engine are settled before any topology is drawn.
8.4
Scale-Out Fabric: Protocols, Standards & Transport
A scale-out protocol is a 3–5 year bet on who supplies your switches, how much link rate survives as goodput, and whether you can leave the vendor whose firmware you depend on.
8.5
Scale-Out Topology, Sizing & Oversubscription
Oversubscription and topology turn a GPU count into a bill of materials, a blocking factor, and an MFU ceiling; the 2026 error is sizing a training fabric for inference, or the reverse.
8.6
Congestion Control, Load Balancing & In-Network Compute
A non-blocking Clos only delivers its bandwidth if congestion control, load balancing, and in-network reduction all work; lose any one and the fabric runs near 50% efficiency.
8.7
Management, Out-of-Band Fabric & PTP/IEEE-1588 Timing
An AI cluster needs an out-of-band network to reach a wedged node when the data plane is dead, and a PTP timing plane to put every log on one clock.
8.8
Scale-Across: Multi-Campus & Cross-Region Fabric (DCI for Distributed Training)
When no single campus can be energized fast enough, the question becomes how to split a synchronous job across buildings — a choice that reshapes optics, transport, training algorithm, and failure model together.
8.9
Physical-Layer & Interconnect Taxonomy
Every link bets on the cheapest, lowest-power medium that closes the budget at the required bit-error rate; at 224G-per-lane that bet is lost a metre sooner than a SerDes generation ago.
8.10
CPO, Fiber Plant & Structured Cabling
Fiber outlives the optics it carries, so pull single-mode glass once for two generations ahead; the live fork is whether the laser stays on the faceplate or moves onto the switch package.