The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 8.1

In this chapter · 6 sections
Term help

Network Fundamentals & AI Traffic Characterization

In a synchronous AI cluster the network sets the pace of every accelerator: the job runs at the speed of its slowest collective and stalls on its longest tail.

GOODPUTPOWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. Where you draw the scale-up / scale-out / scale-across boundaries — because that partition decides which traffic rides a memory-semantic fabric such as Blackwell NVLink 5 (1.8 TB/s per GPU bidirectional aggregate), which rides a 400–800 Gb/s-per-direction packet interface, and which crosses a WAN. Use the dated platform records in this chapter and compare the same direction before assigning workload bytes and deadlines; the rest of Part 8 inherits that choice.
  2. Which parallelism strategy (TP/EP vs PP/DP) maps onto which network tier — get the mapping wrong and an exposed all-reduce sized for the dated Blackwell NVLink 5 profile (1.8 TB/s bidirectional per GPU) lands on a 400 Gb/s-per-direction NIC and leaves model FLOPs waiting on communication. Keep it local when the crossing misses budget; cross deliberately when memory capacity or a tested schedule justifies the cost.
  3. Whether your design target is raw link rate or delivered goodput — the two diverge under tail latency, congestion, and the BSP barrier, and only one of them shows up in the training run's wall-clock time.
  4. Whether you are building a training fabric (synchronous steps, bandwidth and tail delay, job-completion-time) or an inference fabric (request SLOs, tokens-per-second-per-dollar) — independent serving replicas can be loosely coupled, while distributed TP/EP, MoE dispatch and KV transfers put synchronization inside each request. Those dependencies set bisection and oversubscription.
  5. Which traffic-characterization artifacts (collective profile, message-size histogram, incast/tail budget) must exist before you size switches, optics, and buffers — because a fabric sized from peak link rate instead of measured traffic is over-built where it does not matter and under-built where it does.
Illustrative — stated assumptions. Place each communication group against its own byte, dependency and overlap budget. The chapter’s training case compares healthy operation with a lost uplink; loss of a whole spine is a separate cut test in Chapter 8.5. A local group can participate in hierarchical exchange without treating the whole job as one domain. Reserve and test an inter-site path only when placement requires it.

In a single-GPU world the network is an afterthought — feed data in, ship results out. In a frontier AI cluster the interconnect is a first-class component of the computer, and for the largest synchronous jobs it gates all the others. A modern training run is one program spread across tens of thousands of accelerators that must agree, repeatedly and synchronously, on the value of a shared set of gradients. That agreement is a collective communication — an all-reduce, all-gather, or reduce-scatter — and until it completes, every GPU that finished its share of the work sits idle at a barrier waiting for the slowest participant. So the network does not carry bytes between the computations; it sequences the computation itself. Misjudge it and the expensive accelerators run at a fraction of their FLOPS, burning wall-clock while they wait on each other.

The rest of Part 8 builds on the primitives set here: the three-network model (scale-up, scale-out, scale-across) and why its internal boundaries are moving; the collective primitives and how parallelism strategies map onto each network tier; the tyranny of the tail and the BSP barrier semantics that turn a one-in-a-thousand slow link into a fleet-wide stall; the divergence between raw link rate and delivered goodput that decides whether a fabric earns its capital; the two traffic profiles — training versus inference — and job-completion-time as the north-star metric that reframes every downstream networking decision.

The three-network model

An AI cluster is not one network. It is a hierarchy of at least three, each with a different bandwidth, latency, reach, and cost-per-bit, and the single most consequential network-design decision is where you draw the boundaries between them. The hierarchy spans roughly three orders of magnitude in per-GPU bandwidth and four in latency, and traffic that lands on the wrong tier is either ruinously expensive or ruinously slow.

Scale-up is the memory-semantic fabric inside a tightly-coupled domain — historically a node, now including the 72-GPU GB200 NVL72 rack and NVIDIA’s March 2026 forecast of a 576-GPU Rubin Ultra domain across eight racks, each with its own topology and software contract. It is the fastest and most expensive tier: NVLink 5 delivers ~1.8 TB/s per GPU bidirectional (~0.9 TB/s per direction), an NVL72 rack carries ~130 TB/s of aggregate injection bandwidth, and supported peer load/store operations let a GPU read or write another GPU’s memory without a host-staged message; the platform’s ordering rules still apply, and remote access does not imply local latency or cache coherence. This is where you want tensor-parallel and expert-parallel traffic to live, because those collectives are bandwidth-hungry on every layer. → Chapter 8.2.

Scale-out is the packet-switched fabric that stitches domains into a cluster of tens of thousands of GPUs. It is roughly an order of magnitude slower per GPU — a 400 Gb/s-per-direction NIC (50 GB/s) is ~1/18th of NVLink-5's ~0.9 TB/s-per-direction figure; 800 Gb/s per direction (100 GB/s) makes the ratio ~9:1 before overhead — and an order of magnitude cheaper. The semantics are message-passing (RDMA over InfiniBand or RoCE/Ethernet), and this is the natural home for data-parallel gradient all-reduce and pipeline-parallel point-to-point. → Chapter 8.4 and Chapter 8.5.

Scale-across is the newest tier and a creature of the power-bound era: the WAN/DCI fabric that couples multiple campuses when no single site can be energized at the scale a frontier run demands. Crossing a site boundary does not select the algorithm; measured RTT, bandwidth, and the job's communication pattern do. Qualified cross-campus or intra-metro paths can preserve synchronous methods, while longer regional paths move the design toward hierarchical or relaxed-synchrony methods and compression when the bandwidth budget demands it. Power scarcity can push a frontier lab toward multiple campuses when one site cannot supply the required compute, but the protected-route budget must still admit each workload; the grid constraint does not make synchronous traffic fit a WAN. → Chapter 8.8; the power-bound rationale in Chapter 16.1.

Around these three sits a fourth network that every deployment has and no fabric budget should forget: the front-end network — the CPU-side Ethernet carrying data ingestion, checkpoint traffic, storage access, user/API serving, and cluster services. Meta describes its training clusters as running two independent networks — a front-end for ingestion, checkpointing, and logging; a back-end for training — and NVIDIA's reference architectures split the same functions into distinct storage and in-band management fabrics beside the compute fabric (Meta, SIGCOMM 2024; NVIDIA reference architectures, 2025). It is ordinary Ethernet in kind but not in size, and its sizing is a live decision with an order of magnitude inside it. The front-end rate follows concurrent ingestion, checkpoint save/restore, KV movement and service traffic, together with their deadlines. Count the NIC ports actually attached to storage and the uplinks shared with management; dividing a tray aggregate by GPU count does not establish workload demand. A reference configuration is useful only with its host connections, storage-device connections and oversubscription boundary intact. The NVIDIA GB200 reference architecture describes those as a separate fabric. Front-end sizing lives with Chapter 9.2 and Chapter 9.3, and management recovery with Chapter 8.7.

Collective primitives and the parallelism mapping

Distributed training is built from a small vocabulary of collective communication primitives, and almost all the traffic in a cluster is one of them. All-reduce (sum a tensor across all ranks, return the result to all) is the workhorse of data-parallel gradient synchronization. All-gather and reduce-scatter (the two halves of which a ring all-reduce is composed) move sharded weights and activations under tensor and sharded-data parallelism. All-to-all is the signature of mixture-of-experts: each token is routed to its selected experts; the byte matrix can be sparse and skewed, stressing a destination or a cut rather than loading every path uniformly. Point-to-point sends move activations forward and gradients backward across a pipeline stage.

The engineering content is the mapping: which parallelism dimension generates which collective, and which tier can carry its exposed exchanges within budget. Get the mapping right and each collective lands on a fabric sized for it; get it wrong and you push a bandwidth-bound collective onto a tier that cannot feed it, and the job stalls. The four parallelism dimensions are nested precisely so they can be mapped onto the bandwidth hierarchy.

Parallelism → collective → network-tier mapping
ParallelismDominant collectiveFrequencyTarget tierWhy it lands there
Tensor (TP)All-gather / reduce-scatter (per layer)Every layer, every stepScale-up (NVLink)Keep local when per-layer exposed exchanges exceed the scale-out budget; test cross-domain TP if capacity requires it
Expert (EP, MoE)All-to-all (token routing)Every MoE layerScale-up; spills to scale-out as EP degree growsBisection-stressing; wider scale-up domains raise the EP ceiling before it spills
Pipeline (PP)Point-to-point (activations / gradients)Per micro-batch boundaryScale-out (intra-cluster)Sparse transfers; small microbatches expose latency. Scale-out fits when the schedule hides or budgets the transfer
Data parallel / FSDPDP commonly all-reduces gradients; FSDP can all-gather parameters before wrapped modules and reduce-scatter gradients after backwardDepends on wrapping, prefetch, bucket, model and runtimeUsually scale-out; prove any scale-across method separatelyMeasure communication timing and overlap before selecting oversubscription or WAN placement
The placement rule follows the workload deadline: bandwidth-hungry, every-step collectives stay on scale-up when a crossing would delay compute; less frequent or latency-tolerant traffic spills to scale-out and beyond when overlap or the deadline absorbs it. Bandwidth figures are 2025–2026 NVIDIA platform reference points; the selected configuration sets the comparison, with sources and dates in keynumbers.

The table is a placement hypothesis to test. The design move is to keep the leftmost, highest-frequency rows (TP, EP) inside the scale-up domain where the local NVLink fabric’s bandwidth can keep the collective off the critical path when the selected schedule overlaps its completion, and to let the lower-frequency rows (PP, DP) spill onto the cheaper scale-out fabric when their measured timing and overlap tolerate lower bandwidth or oversubscription. This is why domain size dominates the fabric design: each GPU added to the domain can keep another chunk of TP/EP traffic off the slower fabric when the runtime places that state and exchange locally. It is also why mixture-of-experts inference reshaped fabric design — wide expert parallelism benefits when its exposed dispatch/combine traffic stays on NVLink, but enough expert capacity and a tested hierarchical NVLink/RDMA schedule can justify crossing domains. → MoE inference and EP degree in Chapter 8.2.

Tail-latency tyranny and the BSP barrier

One fact separates AI networking from every other kind of data-center networking: a synchronous training step is a bulk-synchronous-parallel (BSP) computation, and a BSP barrier runs at the speed of its slowest participant. When ten thousand GPUs all-reduce a gradient, the collective is not done when the median link finishes — it is done when the last one does. Average bandwidth is almost irrelevant; what governs wall-clock is the tail. A fabric with superb median latency and a fat 99.9th-percentile tail will deliver worse training throughput than a fabric with mediocre median latency and a tight tail. This is the tyranny of the tail, and it inverts the intuition every web-scale network engineer brings to the problem.

The traffic pattern depends on the collective algorithm: parameter-server aggregation and some tree/reduction fan-ins produce synchronized many-to-one incast, while ring all-reduce produces synchronized neighbor flows with a long dependency chain, so its slowest transfer stretches the collective tail. The selected algorithm's traffic matrix determines whether switch buffers, flow control (PFC), or packet loss drive that tail. One slow link, one congested switch, one mis-tuned congestion-control loop, and the barrier that should have closed in microseconds drags for milliseconds, multiplied across every step of a months-long run. Circulating industry write-ups put percentages on this effect, but the concrete failure mechanism is a training barrier stretched by network delay: in a 175B-parameter job, a congested queue or a retransmission can leave ready GPUs waiting for the last rank. Those percentages cannot size a fabric without the model, GPU count, collective algorithm, loss process, runtime and step-time denominator that produced them; a theoretical packet-loss result is a sensitivity to test, not a measured loss of compute on the buyer's cluster. The job's own completion trace decides whether removing that delay recovers enough accelerator time to justify the fabric change. What transfers between clusters is the mechanism: barrier-coupled work turns a fabric defect into fleet-wide idle.

A 400 Gb/s port does not deliver 400 Gb/s of useful gradient synchronization: protocol overhead, congestion, retries and barrier wait stand between the rate on the invoice and the training run’s wall-clock. A faster port earns its price when useful communication advances the job sooner. MAC link rate is the Ethernet service rate, distinct from coded SerDes signaling. Payload throughput is useful delivered bytes divided by transfer time in a named direction; payload efficiency divides that rate by the corresponding MAC byte rate. Collective algorithm bandwidth is the stated message size divided by operation time; NCCL bus-bandwidth normalization depends on operation and ranks and is not switch telemetry. MFU divides useful model FLOPs by elapsed seconds times participating accelerators’ declared peak FLOPs/s at the selected precision. Exposed communication is dependency delay after useful overlap, divided by step time when reported as a fraction. Job goodput counts useful completed work per wall-clock time across the declared checkpoint/recovery window; that operations ledger lives in Chapter 14.1.

The protocol choice is the cleanest illustration. InfiniBand's effective throughput and latency belong to the tested adapter, switch, collective and message size; a switch-hop latency cannot stand in for application completion. Untuned RoCEv2 over Ethernet can show materially higher and more variable latency with a long tail when congestion or reordering triggers the selected NIC's retry behavior; go-back-N receivers repeat traffic that a receiver with supported selective retry or out-of-order placement can avoid. NVIDIA's 2024-10-28 Colossus announcement reports roughly 95% data throughput for Spectrum-X while training Grok. That is one vendor-reported workload and configuration, not a universal Spectrum-X result or evidence for Ultra Ethernet deployments. Meta's 2024-08-05 production account describes the topology, routing and congestion tuning that made RoCE viable for its distributed training. The lesson is that tuning and topology earn their cost when they shorten the same job's exposed communication; a faster link earns its cost when that job is still bandwidth-limited. The NCCL-tests definitions keep collective comparisons on the same denominator. → transport qualification in Chapter 8.4.

For n=64, each rank sends W=2(n−1)S/n=2×63/64×4 GiB=7.875 GiB=8,455,716,864 bytes. Healthy communication is W/(45×10⁹ bytes/s); step time is 800 ms + max(0, W/(45×10⁹)×1,000−150 ms), or 840 ms at 10 ms precision: pass. One lost uplink leaves R=45×31/32 GB/s; the same expression gives 840 ms at 10 ms precision: pass. Unrounded times differ although displayed results coincide. Keep TP local and admit DP only after the built fabric sustains the assumed tail service; this step alone does not justify a faster NIC.

The flip is growth on a fixed surviving cut. With n/2 senders, R(n)=31×45×2/n GB/s and communication is 2(n−1)×4×2³⁰/(2×31×45×10⁹) seconds. It must be no more than (1,100−800+150) ms. The crossover is about 147 ranks: the largest balanced even group that passes is 146; 148 fails. Extra ranks need extra leaf ports: this sensitivity does not enlarge the original two leaves. Fewer bytes or more overlap moves the crossover outward. The NCCL-tests method supplies normalization; Chapter 8.5 tests spine loss and counts expansion.

840 ms / 840 msderived
Healthy / one lost link: no displayed change at 10 ms precision
Rounding to 10 ms hides a small slowdown; both states meet their stated deadlines.
Scope & caveats

64 DP ranks/rail; 4 GiB/rank; 45 GB/s healthy, 31/32 service after one uplink loss; 800 ms work, 150 ms overlap. Rounded to 10 ms; spine loss is separate in Chapter 8.5.

1.8 TB/s bidirectional (~0.9 TB/s per direction)
NVLink-5 per-GPU scale-up bandwidth, bidirectional (NVL72 = ~130 TB/s rack aggregate); ~18x a 400 Gb/s scale-out NIC per direction (900 vs 50 GB/s each way)
Scope & caveats

Named NVIDIA generation/platform, bidirectional aggregate convention; not delivered collective bandwidth, cache coherence or arbitrary mixed-platform performance.

1.07 µs
HDR InfiniBand: 8-byte host-memory MPI ping-pong mean, EPYC Rome / ConnectX-6 (HPC Advisory Council, July 2020)
Scope & caveats

Historical 8-byte host-memory MPI ping-pong mean, not p99, GPU-memory latency or a RoCE comparison. Daytona_X / EPYC Rome / ConnectX-6 HDR; OSU 5.6.2, HPC-X 2.7.0, OFED 5.0.2; local core 80, 10,000 timed/warm-up iterations.

33% → 15%estimate
unreproduced industry figure for GPU idle-time reduction, congested vs well-engineered fabric — no published cluster, model or denominator
Scope & caveats

Unreproduced industry figure. The source is a vendor blog post that states the range without a cluster, model, collective algorithm, denominator or measurement method; no primary experiment has been located.

Not transferable to a specific cluster — use it as the shape of the effect and measure your own idle fraction against your own step time.

1:1 vs 2:1–3:1
published scale-out examples (1:1, 2:1–3:1, reported 7:1) — workload-specific observations, not defaults; derive from measured traffic, topology and SLO
Scope & caveats

named reference designs and deployments with different traffic, topology, placement and service objectives

Derive oversubscription from the measured traffic matrix, collective/request mix, topology, failure headroom and SLO; validate it on the target fabric.

8.4% (10.0% incl. NIC)
share of Llama-3 training interruptions attributed to network switch/cable faults
Scope & caveats

Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

Training vs inference traffic profiles

Training and inference often produce different traffic, but the label does not select the fabric. Synchronous collectives are bursty and bisection-hungry, so measured step-time and MFU targets often justify high bisection. Many local inference requests permit upper-tier multiplexing, while distributed MoE, KV movement, and prefill/decode disaggregation can require substantial scale-out capacity. Published fabrics span 1:1, 2:1–3:1, and reported 7:1 examples. Chapter 8.5 turns those observations into a per-tier ratio through its locality, exposure and failure-headroom tests; this chapter supplies the traffic characterization those tests read.

The 2025–2026 twist is that inference traffic is no longer trivial. Reasoning models emit long decode sequences, mixture-of-experts routing creates real all-to-all traffic, and prefill/decode disaggregation moves the KV-cache across the fabric — so inference, once a single-node afterthought, now exercises the scale-up domain hard. The characterization still differs from training (request-SLO-bound, with synchronization inside distributed kernels), but the gap is narrowing, and a fabric scoped for last-generation single-node inference can be undersized for MoE reasoning serving. → inference fabric demands in Chapter 1.3; MoE all-to-all in Chapter 8.2.

Training vs inference: how the two workloads characterize traffic
PropertyTraining fabricInference fabric
CouplingTight — synchronous BSP barrier every stepIndependent replicas can be loosely coupled; distributed TP/EP and KV transfers impose request dependencies across the selected domain
Traffic shapeBursty east-west; silent-then-flood at the collectiveRequest arrivals, expert dispatch and KV transfers; batching can synchronize peaks
Governing metricJob-completion-time (goodput / MFU)Latency SLO (TTFT, TPOT) at target throughput
Tail toleranceBounded by the step deadline at each exposed dependencyBounded — tail shows up as p99 latency, not a global stall
OversubscriptionDerive from collective traces + step-time/MFU targetDerive from request/KV/EP traces + tail-latency SLO
Siting consequencePower-first; can be far from usersLatency-first; geo-distributed near users
Both can use either protocol family. These are common traffic tendencies; measured traffic, placement, topology, failure headroom, and SLO — not the workload label — set each tier's blocking ratio.

The transfer allowance is 250−180=70 ms. Eight requests carry 8×1 GiB=8,589,934,592 bytes. Across racks, completion is 180 ms +8,589,934,592/(45×10⁹)×1,000 ms, about 370 ms: fail. With local service of 300 GB/s it is about 210 ms: pass, rounded to 10 ms. Select local prefill/decode placement, accepting its memory and scheduling constraints to avoid the request-tail miss.

Batch size reverses the choice: B×2³⁰/(45×10⁹) ≤0.070 gives B≤2.93. Two concurrent transfers give about 230 ms: pass; three give about 250 ms but fail the unrounded limit. Select cross-rack placement at two transfers or fewer, or acquire more service capacity. Hand Chapter 9.7 the KV representation, bytes/request, burst, ownership-transfer event and time allowance; it owns KV sizing. The Dynamo disaggregated-serving design supplies the mechanism. Validate the joint request tail in Chapter 13.7.

Job-completion-time as the north-star

Every networking decision in Part 8 ultimately answers to one metric, and it is not bandwidth, latency, or port count — it is job-completion-time (for training) and its inference twin, tokens-per-second-per-dollar at the SLO. JCT is the wall-clock from job start to a trained model, and it folds in everything the spec sheets hide: the tail, the congestion, the goodput, the failures and restarts. A fabric is good if and only if it lowers JCT per dollar of capital and per watt of power — and a fabric can have best-in-class link rate, best-in-class median latency, and still lose on JCT because its tail or its reliability is poor.

Reframing networking around JCT changes what you optimize. You stop chasing peak bisection bandwidth as a vanity number and start chasing the combination that minimizes wall-clock: tight tail, high goodput, fast recovery from the inevitable link and switch failures (which the Llama-3 run attributed ~8.4% of interruptions to). It also re-prioritizes the whole part — congestion control, in-network reduction, adaptive routing, and topology matter because they move JCT, not because they look impressive in a bake-off. This is the same shift from availability to goodput that the reliability chapter makes for the facility: the question is never 'is the link up?' but 'how much useful training did the fabric deliver this week?'. → goodput vs availability in Chapter 12.2; checkpoint/restart math behind failure recovery in Chapter 9.4.

Deep dive: why the scale-up / scale-out boundary keeps moving

The three-network model is not a fixed partition — its internal boundaries migrate with each hardware generation, and that movement is one of the most important dynamics in AI infrastructure. The pressure is bidirectional. Pushing the scale-up boundary outward: NVIDIA's 2024-10-15 platform account contrasts eight GPUs in HGX H200 with 72 in GB200 NVL72, while its 2026-03-16 design account forecasts 576 GPUs across eight racks for Vera Rubin Ultra NVL576 — because keeping exposed TP and EP exchanges on the local NVLink memory-semantic fabric instead of spilling them onto NICs whose delivered bandwidth misses the exchange budget is a lever on model-FLOPS-utilization for large dense and MoE models when the displaced exchanges were exposed on the critical path. Bigger domains raise the tensor- and expert-parallel ceilings and widen the all-to-all that MoE inference depends on.

Pulling it back: the copper reach wall. NVLink over copper is bounded by the qualified package, board, connector and cable route; a domain that outgrows that budget needs a shorter layout, active copper or supported optics. Optical engines add power and replaceable parts, while moving them onto a package changes which failure requires a module, board or switch repair. So the scale-up boundary sits exactly where the bandwidth benefit of a larger domain still outweighs the cost, power, and blast-radius penalty of crossing into optics. Meanwhile a brand-new boundary opened at the top — scale-across — because the grid cannot energize a single site large enough for a frontier run, pushing the program toward multi-campus training over a WAN whose protected routes must first admit the workload’s bytes and synchronization deadlines. The practical consequence for a facility designer: provision for the boundary to move. A hall plumbed and powered only for today's domain size strands the option to adopt the next generation's larger one. → copper-vs-optics in the domain in Chapter 8.2; the roadmap in Chapter 16.2; scale-across in Chapter 8.8.

Deep dive: characterizing traffic before you size the fabric

The recurring mistake is sizing a fabric from peak link rate and rack count instead of from measured traffic — which over-builds where it does not matter and under-builds where it does. A defensible fabric design starts from three traffic-characterization artifacts, derived from the workload's collective profile rather than a vendor reference diagram.

  • Collective profile. Which collectives the job runs (all-reduce, all-gather, all-to-all, point-to-point), at what frequency, and on which parallelism dimension — because that is what tells you how much traffic wants scale-up versus scale-out, and therefore the domain size and the blocking factor. An MoE model with heavy all-to-all and a dense model doing ring all-reduce produce completely different fabric demands at the same GPU count.
  • Message-size histogram. The distribution of message sizes, because small-message latency and large-message bandwidth are served by different switch and protocol choices, and a fabric tuned for one is mediocre at the other. Gradient all-reduce is large-message bandwidth-bound; pipeline activations and control traffic are latency-bound.
  • Incast and tail budget. The expected incast ratio at the collective and the p99/p99.9 tail-latency budget the BSP barrier can tolerate — because that, not average bandwidth, sizes buffers, picks the congestion-control regime, and decides whether you need deep-buffer switches, adaptive routing, or in-network reduction.

These artifacts are the network analogue of the workload profile sheet that scopes the whole facility. Produce them before you count switches and optics, and the sizing in Chapter 8.5 becomes a derivation rather than a guess.

This chapter sets the foundation the rest of Part 8 builds on. The scale-up fabric and the moving domain boundary are engineered in Chapter 8.2; the switch, NIC, and DPU silicon in Chapter 8.3; the scale-out protocol war (InfiniBand vs RoCE vs Spectrum-X vs Ultra Ethernet) in Chapter 8.4; topology, sizing, and the oversubscription decision in Chapter 8.5; congestion control, adaptive routing, and in-network compute — the tail-fighting toolkit — in Chapter 8.6; and scale-across multi-campus training in Chapter 8.8. The training/inference fork that this chapter reads from the network is scoped facility-wide in Chapter 1.2 and Chapter 1.3; the goodput-over-availability reframing in Chapter 12.2; the failure-and-restart math behind JCT in Chapter 9.4; and the metric definitions (goodput, MFU, JCT, $/M-tokens) in Chapter 0.3.

Choose the tier boundary from the exchanges that remain exposed, then carry healthy and degraded deadlines into the fabric order. Keep a group local when crossing fails that budget; split it when memory capacity demands the crossing and the measured schedule can afford it. Choosing on port rate alone buys capacity the job may never use while leaving its slowest dependency unpaid for.

Cite this chapter
Fehn, J. (2026). Network Fundamentals & AI Traffic Characterization (Chapter 8.1). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-1-network-fundamentals-and-ai-traffic-characterization (accessed 2026-09-29).
@misc{aidc-8-1,
  author       = {Fehn, Jacob},
  title        = {Network Fundamentals & AI Traffic Characterization (Chapter 8.1)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-1-network-fundamentals-and-ai-traffic-characterization},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit