Chapter 8.6
In this chapter · 6 sections
Congestion Control, Load Balancing & In-Network Compute
A non-blocking Clos only delivers its bandwidth if congestion control and load balancing both work; lose either and an untuned fabric leaves a large fraction of its rated throughput on the floor, while in-network reduction accelerates the collectives that support it rather than gating the fabric.
What you'll decide here
- Whether you run InfiniBand credit flow, RoCE with scoped PFC, or a supported loss-tolerant transport such as UET with its selected trimming configuration — credits and pauses can propagate head-of-line blocking or deadlock, while retries consume time and receiver resources. Prove how each contains an interior hotspot, destination overload and stalled receiver.
- Whether you accept flow-level ECMP and its hash-collision tax on a handful of fat elephant flows, or use extra QP entropy, weighted/resilient ECMP, flowlets, adaptive routing or per-packet spray. Match movement to traffic gaps and pay for the receiver state and reordering it needs; more paths cannot raise the destination’s service rate.
- Whether you offload collectives into the switch (SHARP, NVLink-SHARP) to reduce endpoint traffic and reclaim GPU SM work when the operation, dtype, size and group concurrency qualify — and accept the vendor, radix and topology constraints only if endpoint fallback still meets the deadline after an aggregation resource or path is lost.
- How aggressively to tune the DCQCN parameter space (ECN marking thresholds, CNP cadence, PFC headroom) versus buying a fabric where the vendor has pre-tuned it — because an untuned RoCE fabric runs at a fraction of its rated throughput.
- What congestion telemetry you instrument now (per-queue depth, ECN/CNP counters, PFC pause duration, RNR/retransmit) so the operations team in Part 10 and Part 14 can find the one straggler link before it taxes every step.
Chapter 8.5 sized a topology that, on paper, is non-blocking — enough bisection bandwidth that every GPU can talk to every other GPU at line rate. That paper number is a lie until you solve three problems the topology alone does not. AI traffic is the pathological worst case for a packet network: a handful of enormous, long-lived elephant flows per host, all synchronized to the same collective, all hitting the wire at the same microsecond, all converging on the same destinations during an all-reduce. A synchronized collective can hit the fabric as a coordinated stampede followed by an idle gap; do not rely on background mice to smooth that burst. Mixed inference and storage traffic add their own demand to the same queue and admission budget. In that regime, congestion control, load balancing, and in-network compute stop being tuning knobs and become the difference between the bandwidth you paid for and the bandwidth you get.
The chapter works through the lossless-vs-lossy fork (and the PFC pathologies that haunt the lossless path), the DCQCN parameter space you must tune or buy pre-tuned, the load-balancing fork between flow-hashed ECMP and adaptive/per-packet spray (and the reordering bill the latter sends to the NIC), and in-network compute — SHARP and friends — a lever that can improve goodput and energy per step when an eligible reduction removes critical endpoint work and traffic at lower whole-system energy cost. It closes on the telemetry that makes all of it observable, because a congestion problem you cannot see is a goodput problem you cannot fix. → fundamentals and the goodput framing in Chapter 8.1; the protocol/transport choices these mechanisms ride on in Chapter 8.4.
Why AI traffic breaks ordinary congestion control
Datacenter congestion control was designed for the web-scale workload: millions of short flows, bursty but uncorrelated, where TCP's loss-and-backoff and a little ECN keep buffers shallow and tails short. AI training inverts every one of those assumptions. A single all-reduce on a 100k-GPU cluster is one logical operation decomposed into thousands of simultaneous, multi-gigabyte point-to-point transfers that all start on the same barrier and must all finish before the next step can begin. The collective runs at the speed of its slowest flow — the tail-latency tyranny of Chapter 8.1 made concrete at the packet layer. One congested link, one hash collision, one PFC pause that ripples the wrong way, and the straggler it creates stalls every GPU in the job on the next bulk-synchronous barrier.
Three failure modes follow, and the rest of the chapter defends against each in turn. Buffer pressure and loss from algorithm-specific traffic: many-to-one incast in parameter-server and tree fan-ins, or synchronized neighbor flows in ring all-reduce, which the congestion-control loop must absorb. Path imbalance — a few elephant flows colliding on one physical link while parallel links sit idle — which load balancing must spread. And endpoint overhead — reduction arithmetic on the GPU and bytes moving through memory, PCIe and the NIC — which eligible in-network operations can reduce; the host attachment and library path determine which transfers and control work remain. Each carries a distinct downstream cost in goodput, power, or lock-in.
The lossless-vs-lossy fork and the PFC trap
RDMA — the zero-copy, kernel-bypass transport that makes GPU-to-GPU networking fast — was born assuming a lossless fabric. InfiniBand delivers losslessness natively with credit-based flow control: a sender never transmits unless the receiver has advertised buffer credits, so packets are never dropped for lack of room. A RoCEv2 lossless deployment adds a separate Ethernet mechanism: Priority Flow Control (PFC), a per-priority PAUSE frame that tells the upstream link to stop sending before a buffer overflows. PFC works — and it is the source of the nastiest pathologies in AI networking.
The first is head-of-line blocking. PFC pauses an entire traffic class, not a flow. When a downstream buffer fills, the PAUSE stops every flow in that priority on that link, including flows whose destinations are perfectly idle — innocent bystanders, the "victim flows." The second is the PFC pause storm and, in the worst case, deadlock: because PAUSE is hop-by-hop back-pressure, a congestion hotspot propagates upstream toward the sources, and in a topology with a cyclic buffer dependency (which fat-trees can form under certain failure or routing conditions) the pauses can form a permanent standstill where no switch can drain because every switch is waiting on another. A deadlocked fabric does not degrade — it stops. The defense is a PFC watchdog: a timer on every queue that, if a port stays paused beyond a threshold, drops the stuck traffic and logs it rather than letting the deadlock persist — trading a localized packet loss for fabric-wide survival.
That is why the loss model belongs in the system contract, alongside loss recovery and receiver resources. Meta’s 2024 production account describes its PFC-based 400G back-end, with end-to-end congestion control disabled in that reported deployment; its controlled topology and traffic are part of that result. Ultra Ethernet specifies congestion signaling and loss recovery, with trimming available in the selected switch/transport configuration: the switch discards payload while forwarding a shortened packet that identifies the discarded payload, allowing supported selective recovery to retry the missing data rather than the later window. Feedback still travels a real route and can itself be delayed. Choose the transport’s supported combination, then prove the fault response rather than inferring it from the words lossless or lossy. → operation and completion semantics in Chapter 8.4; switch buffer resources in Chapter 8.3.
DCQCN tuning: match the loop to the bottleneck, or buy a supported profile
DCQCN — Data Center Quantized Congestion Notification — is the end-to-end loop that keeps a RoCE fabric out of PFC pause and away from loss. The mechanism is elegant: a switch experiencing early congestion marks packets with ECN (sets the Congestion Experienced bit) before its buffer is full; the receiving NIC notices the mark and fires a Congestion Notification Packet (CNP) back to the sender; the sending NIC quantizes its injection rate down in response, then probes back up as the congestion clears. Done right, DCQCN holds buffer occupancy in a sweet spot — full enough to keep links busy, empty enough to never trigger PFC — and PFC becomes the backstop it was meant to be rather than the daily driver.
Done wrong, the same feedback can oscillate between excess queueing and underused links. DCQCN exposes a large, coupled parameter space — the ECN marking thresholds (Kmin and Kmax: when to start marking, when to mark every packet), the marking probability slope, the CNP generation cadence, the rate-increase and rate-decrease step sizes and timers, all interacting with the per-queue PFC thresholds and headroom. Set the ECN threshold too high and you mark too late, the buffer overruns, PFC fires, and you are back in pause-storm territory. Set it too low and you throttle senders that were not actually congested, leaving bandwidth on the table and inflating job completion time. The parameters are also topology-, speed-, and workload-dependent: a move from 400G on a two-tier Clos to 800G on a three-tier one changes in-flight bytes and feedback delay, so the old thresholds need requalification. An untuned RoCE fabric can leave rated capacity unused, but a quieter queue obtained by over-throttling senders is not recovered throughput: measure the job’s step tail, receiver drain and queue trajectory together.
First locate the service constraint. An interior egress queue rising while alternative paths are idle selects a path-distribution remedy. Every path converging on a full destination link selects admission or more destination capacity. A quiet wire with depleted posted receives selects receive-resource repair; adequate receives with memory consumption stalled selects the consumer. Meta’s receiver-driven admission account motivates this separation. In UET, NSCC uses network congestion signals to regulate injection, while RCCC uses receiver-granted credits to control arrivals. They address different resource boundaries; protect the returning control traffic and test lost or stale feedback. Direct switch-to-source signaling must have a supported source response and a failure policy. IEEE P802.1Qdw (source flow control) is a PAR-approved project without a published standard, and CSIG remains an expired individual Internet-Draft (draft-ravi-ippm-csig-01, 2024-02-02); product support for either is separate from document status.
Worked decision A: incast backlog and pause headroom
Four inputs offer 4 × 400 = 1,600 Gb/s. With the output draining, excess arrival is (1,600 − 400) × 10⁹ / 8 = 150 × 10⁹ B/s. Backlog after the assumed delay is 150 × 10⁹ B/s × 2 × 10⁻⁶ s = about 0.3 MB, below the 0.40 MB usable allowance. The crossover is 0.40 × 10⁶ B / (150 × 10⁹ B/s), about 2.7 µs. The unrounded threshold is 2⅔ µs; 3 µs produces 0.45 MB and fails. Sustained incast still needs admission.
Ingress PFC headroom is a different ledger. Each paused input needs 400 × 10⁹ bit/s × 2 × 10⁻⁶ s / 8 + 2 × 9,216 B + 32,000 B = about 0.15 MB. The unrounded allocation check is 150,432 B ≤ 160,000 B, with four simultaneous pauses requiring 601,728 B ≤ 640,000 B. At 3 µs, one input needs 200,432 B and fails. Do not deduct egress drain from ingress headroom or double-count shared storage.
Accept buffering only when allocation and measured delay pass both ledgers; otherwise reduce admission or buy more usable buffer. Verify cell rounding, pool competition, pause/resume and watchdog recovery. The Cumulus buffer and RoCE documentation supplies the configuration method; Chapter 8.5 owns the surviving fabric ledger and Chapter 13.7 executes the acceptance test.
| Posture | Loss model | Primary mechanism | Load balancing | Tuning burden | Best fit |
|---|---|---|---|---|---|
| InfiniBand (credit flow) | Buffer-credit flow control; other faults still need recovery | Link credits and supported routing | Mode and endpoint contract | Integrated controls; qualify loaded recovery | Required operations and tails pass on the purchased system |
| RoCE + PFC + DCQCN | Scoped lossless priority; watchdog can deliberately drop | ECN → CNP → sender rate | ECMP or qualified adaptation | Own class, queue, NIC and pool settings | Required operations pass with bounded pause and recovery |
| RoCE + PFC-only | PFC-backed priority | Hop-by-hop back-pressure | Pinned paths or qualified distribution | Requires measured topology and workload control | Select only after the same incast and fault tests |
| Ultra Ethernet (UET) | Selected delivery/loss-recovery profile | NSCC / RCCC and supported trimming | Qualified spray and endpoint delivery handling | Version-matched feature and interoperability evidence | Required profile and completion semantics are implemented |
| MRC (RoCEv2 extension) | Specified trimming and selective acknowledgment | Sender window from ECN and delay | Endpoint-controlled multipath and reordering | Qualify operations, control path and firmware | Workload meets the operation contract in Chapter 8.4 |
If you build a RoCE fabric from merchant switch silicon and standard NICs, you own the DCQCN parameter space — and owning it means a standing investment in network engineers who can read ECN/CNP counters, run incast tests, and re-tune every time you change speed, scale, or topology. The alternative is to buy a fabric where the vendor has done that work and validated it: NVIDIA Spectrum-X bundles switches and SuperNICs with a supported congestion and adaptive-routing profile; NVIDIA’s October 2024 Colossus report describes roughly 95% data throughput for that deployment. A single escalation path can save integration work, but the result does not replace workload acceptance. InfiniBand uses credit flow rather than DCQCN. The decision is a classic build-vs-buy on the network: self-tuned merchant RoCE can trade a lower switch/NIC bid for a recurring tuning liability carried by your engineering team. An integrated fabric converts only its contracted support duties into a line item. Compare acquisition, NIC/switch/optics/firmware support and failure recovery cost at equal service. Pick based on whether you have the network-engineering bench to keep a merchant RoCE fabric tuned through every speed, scale and topology change.
Worked decision B: tune, remove a spine, then recover
Export effective NIC/switch settings and bidirectional counters; the versioned Cumulus profile is the class-map reference, not authority for these proposed thresholds. Verify that background traffic cannot occupy the headroom already reserved. With the stated first-event alpha, the sender rate becomes 400 Gb/s × (1 − 1/16) = 375 Gb/s. Four such senders still offer 1,500 Gb/s to the destination’s 400 Gb/s service; one decrease cannot close sustained incast. Continue feedback and admission until the aggregate arrival rate fits service, or the queue must keep growing.
Replay the same traffic seed and rank placement. Collect healthy tails, queue high-watermarks, ECN/CNP and pause duration. Remove one spine, timestamp withdrawal, verify Chapter 8.5’s surviving ports, and measure transition and converged behavior. Pass requires recovery within 1 s, converged p99 within 1,100 ms, healthy p99 within 900 ms, no corruption and no persistent watchdog-driven pause cycle. Report transition interruption separately. Inject receiver starvation and delayed CNP separately.
HOLD deployment pending the hardware trace. Tune one control family at a time with a rollback snapshot; if marking comes too late, lower the egress thresholds before consuming ingress headroom. If tails rise while queues stay empty, restore the prior rate-recovery profile. The decision flips to accept only when the specified profile passes every bound; a converged p99 above 1,100 ms reverses it even if port utilization looks balanced. Chapter 13.7 records installed acceptance.
Load balancing: ECMP's hash-collision tax vs adaptive routing and packet spray
A fat-tree gives you many equal-cost paths between any two endpoints. The question is how you spread traffic across them, and AI traffic makes the naive answer fail badly. The default is ECMP — Equal-Cost Multi-Path — which hashes each flow's 5-tuple to pick one of the available uplinks. For web traffic with thousands of small flows, the hash spreads load beautifully by the law of large numbers. For AI traffic with a handful of elephant flows per host, the law of large numbers does not apply: two fat flows can hash to the same physical link and collide, saturating it while a parallel link sits half-empty. A collision can persist across the critical collective while those flows keep their hashes, so the straggler remains until a flow ends or the path mapping changes. On a synchronous collective, that one congested link taxes every step. This is the structural reason flow-level ECMP under-delivers on AI fabrics: there are too few flows to hash well, and they last too long to absorb a bad hash.
Two answers exist, and they trade simplicity against the cost of fixing in-order delivery. Adaptive routing lets supported InfiniBand and Ethernet switches steer traffic away from a congested egress toward a quieter path using the congestion information they actually see. The movement unit can be a flow, flowlet or packet; local egress occupancy cannot reveal every downstream bottleneck, and permitted movement must match the receiver’s ordering contract. On merchant Ethernet silicon this has productized as utilization-aware dynamic load balancing: the switch assigns each new flow — or flowlet, a burst separated by an idle gap long enough to re-path without reordering — to the least-loaded member link instead of a blind hash, and RDMA-aware variants recognize individual queue pairs and re-place them when congestion appears. These switch-side mechanisms can recover part of ECMP’s collision tax without replacing the transport, which is attractive in a multi-vendor fabric, provided their packet movement stays inside the receiver’s ordering contract; what they cannot do is let one elephant use several paths at once — that still takes per-packet spray and endpoint reordering. Packet spraying goes further: it abandons per-flow pinning entirely and sprays the packets of a single flow across every available path, allowing a flow to exploit spare capacity on several paths. The cost is that packets now arrive out of order, and classic RoCE's go-back-N retransmission treats out-of-order as loss — the in-order penalty that has historically made spray a non-starter on RoCE. The whole point of Ultra Ethernet's transport is to break that penalty: a qualified UET implementation combines packet-level multipath selection with NIC-side placement, reordering and completion behavior supported by its selected delivery mode, making out-of-order delivery a feature rather than a fault, and exposing per-packet multipathing that was previously locked inside proprietary fabrics. → the in-order penalty and how each transport handles it in Chapter 8.4; the topology that defines how many paths exist to spread across in Chapter 8.5.
Deep dive: why elephant-flow ECMP collisions are worse than they look
The intuition that collisions average out fails when a synchronized collective waits for a long-lived straggler. For the illustrative eight-flow all-reduce over sixteen equal-cost paths, set F = 8 and P = 16; with independently hashed flows, the no-collision probability is P! / ((P−F)! PF) for F ≤ P. A collision matters when the colliding offered rates exceed their shared link’s service; it does not automatically halve each flow’s throughput. If that overloaded flow is the last dependency, one collision can set the whole collective’s pace. Queue depth and per-member utilization distinguish that case from an endpoint that was never sending at line rate. Keep the traffic’s actual QP entropy and lifetime in the calculation, because adding idle paths cannot help a flow pinned to a busy one.
Choose the movement granularity deliberately. Fixed ECMP preserves a stable path; more QPs add entropy but consume endpoint state. Resilient hashing limits disruption when a member changes, while weighted hashing reflects unequal capacities without creating packet-level adaptation. Flowlets require idle gaps exceeding path-delay differences; a continuous elephant supplies no such gap. Adaptive selection uses the information actually exposed by the switch, which can be only its local egress occupancy. Packet spray stripes a flow’s packets over several paths, but the receiver and transport must absorb the resulting reordering; the 400/800G Spectrum-X and UEC endpoint candidates in Chapter 8.3 make NIC placement/completion resources part of that hardware purchase. Test unequal path delay, a lost member and stale downstream feedback before enabling a mode. → implementation support in Chapter 8.3.
| Mode | Information used | Receiver/order gate | Failure to test |
|---|---|---|---|
| Fixed / resilient ECMP | Flow hash; member set | Stable path per flow; qualify remap on failure | Lost member and persistent collision |
| Weighted ECMP / extra QPs | Configured weights / additional hashes | QP resources and transport ordering | Unequal paths and stale weights |
| Flowlet | Idle gap and local load | Gap exceeds differential path delay | Continuous elephant with no useful gap |
| Adaptive | Local queue, or supported downstream signal | Movement granularity matches endpoint support | Stale feedback and route asymmetry |
| Per-packet spray | Per-packet choice over allowed routes | Reorder memory, completion and selective recovery | Reorder-window exhaustion and lost path |
In-network compute: SHARP and collective offload
The previous two sections fight congestion by managing the traffic. In-network compute attacks the problem from the other side: send less traffic in the first place by doing the collective's arithmetic inside the switch. In a conventional all-reduce, every GPU's gradient buffer is shuffled across the network in a ring or tree, summed at each hop on the endpoints, and the result scattered back — the data crosses the fabric multiple times, and every reduction step burns GPU streaming-multiprocessor cycles and PCIe/NIC bandwidth that could have been training. SHARP — NVIDIA's Scalable Hierarchical Aggregation and Reduction Protocol — moves the reduction into the switch ASIC: GPUs send their data up an aggregation tree, the switches sum it in-network as it passes, and only the single reduced result comes back down. Switch reduction can replace repeated endpoint exchanges with aggregation along a tree; count endpoint-facing bytes separately from traffic summed over every fabric hop.
SHARPv4 on the Quantum-X800 generation provides 14.4 TFLOPS of in-network compute, nine times the prior generation, with FP8 support (NVIDIA, March 2024). SHARP and NVLink-SHARP move eligible reduction arithmetic into the switch and free GPU SMs for model work, so compare them with the best eligible endpoint algorithm. For n ranks each reducing S bytes, a bandwidth-oriented ring sends 2(n−1)S/n bytes per rank and receives the same amount. An ideal switch reduction receives S from each rank and returns S to it, so its endpoint-facing ratio to the ring is n/[2(n−1)]. That ratio describes bytes; reducing those bytes can save congestion and energy, but execution time and whole-system power still need measurement. Headers, tree depth, reduction resources, library launch and overlap still cost time. Freeing GPU arithmetic helps only if that work occupies the critical path; a job dominated by another collective or local compute will not inherit a byte-count speedup. NCCL-tests normalization supplies the measurement convention; keep algorithm bandwidth, bus bandwidth and the actual job deadline separate.
UALink Common 2.0, announced in April 2026, specifies in-network compute; SHARP requires supported NVIDIA switch, endpoint and library resources. You cannot bolt that reduction onto a fabric that lacks the engines or aggregation topology. The qualification record therefore names the operation, dtype and accumulation precision, message-size range, group count and concurrency, switch aggregation resources, NIC and library releases, and proof that the library selected offload. Exhaust a reduction resource and remove an aggregation path: record whether the call retries, falls back to the endpoint algorithm or fails the communicator, and whether the fallback still meets its deadline. NVIDIA’s SHARP account supports the mechanism; the UALink family record in Chapter 8.2 separates specification scope from purchasable support. Keep platform eligibility in Chapter 8.2, and the physical cut in Chapter 8.5. Select offload when the measured critical phase and fallback justify the integration and vendor coupling; reverse the choice when resource exhaustion exposes a deadline the endpoint implementation cannot meet.
| Symptom in the cluster | Underlying cause | Wrong fix | Right lever |
|---|---|---|---|
| Fabric-wide slowdown traced to one hot port | Pause propagation or destination overload | Assume every pause is a path collision | Compare destination service, ingress pause and egress queue trajectories |
| One spine link saturated, parallels idle | Hash collision, unequal capacity or stale path state | Enable spray without checking reorder support | Trace flow/QP mapping and surviving members, then choose movement granularity |
| All-reduce slower than link rate implies | Launch, endpoint work, reduction resources or fabric queueing | Assume switch offload fixes every collective | Profile the critical phase and compare eligible offload plus fallback |
| RNR NAKs and retry-driven stalls | Receiver not ready: posted-receive / WQE starvation at the target | Treat it as reordering and disable multipathing | Restore posted-receive / WQE headroom; tune receive queues and consumer rate |
| Cluster healthy but JCT creeping up | Untuned congestion loop leaving BW on the table | Accept it | Re-tune ECN thresholds; instrument telemetry (below) |
Scope & caveats
NVIDIA-reported Colossus deployment result under that workload and configuration, not a universal Ethernet, Spectrum-X, or scheduled/VOQ-fabric figure; Colossus is an adaptive-routing/telemetry RoCE fabric, not the scheduled-fabric category. Meta reports tuning RoCE and InfiniBand GenAI clusters to equivalent performance — no common-workload test crowns either transport.
Scope & caveats
Ideal bytes only, not measured bandwidth or job acceleration; launch, contention, offload eligibility and fallback still determine completion.
Scope & caveats
Vendor headline figure for the platform; realized collective speedup depends on the supported switch, NIC and library combination.
Scope & caveats
Historical initial public UEC 1.0 release. The maintained specification revision is recorded separately; publication alone does not qualify an implementation.
Scope & caveats
Historical 8-byte host-memory MPI ping-pong mean, not p99, GPU-memory latency or a RoCE comparison. Daytona_X / EPYC Rome / ConnectX-6 HDR; OSU 5.6.2, HPC-X 2.7.0, OFED 5.0.2; local core 80, 10,000 timed/warm-up iterations.
Scope & caveats
Analyst estimate for 100k H100-class RoCEv2 fabrics; not a measured result on a named workload, and not portable to a different topology, speed or collective mix.
Scope & caveats
Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.
Congestion telemetry: you cannot fix what you cannot see
Every mechanism above produces evidence, and a fabric can leave expensive links idle while the operations team watches the wrong counter. Use queue, retry, pause and receiver evidence at a boundary and time resolution that distinguish competing causes. The congestion-control loop emits ECN-marked packet counts and CNP rates — a rising CNP rate on a link is the earliest signal that DCQCN is being asked to throttle, and a leading indicator of a hotspot before it becomes a PFC pause. PFC itself emits pause-frame counts and cumulative pause duration per priority per port — any sustained pause time is a red flag, and pause storms show up here long before they deadlock. The transport emits retransmit, RNR-NAK, and out-of-order counters that distinguish a real loss problem from a load-balancing reordering problem. And the switches expose per-queue buffer occupancy and microburst histograms that show whether you are riding the DCQCN sweet spot or skating the edge of overflow.
The hard part is correlation, not collection. A straggler flow shows up as elevated JCT at the framework, a congested queue at one switch, a CNP spike at one NIC, and a pause counter three hops upstream — and tying those to the one link that needs attention requires time-aligned telemetry across the whole fabric, which is why precise time synchronization (PTP/IEEE-1588, → Chapter 8.7) sets the uncertainty of cross-device event ordering; it cannot reconstruct a burst that was never captured. Streaming these counters, baselining them, and alerting on deviation belongs to the observability stack: fleet observability and GPU/network health in Chapter 10.6, and DCIM-grade facility-plus-fabric telemetry correlation in Chapter 14.2. Instrument the counters before first light — the first large training run is the worst time to discover you have no visibility into why a step is slow.
Join job and rank → NIC/QP → priority and queue → physical port/lane → route member. Observe whether a congested egress first precedes ECN marks, receiver CNP and sender-rate decline; concurrent growth on every route into the same destination instead selects receiver admission. Re-run with the consumer deliberately slowed: rising receive-resource stalls while the wire is quiet falsifies the path-congestion diagnosis. Preserve clock origin and uncertainty, sample period, histogram window, resets, wraparound and missing/coalesced samples. Polling can reveal trends; queue high-watermarks and triggered hardware captures are needed for bursts shorter than the polling interval. Streaming does not create finer source timestamps. gNMI transports updates, IOAM defines packet-carried fields, and implementation-specific INT/IFA modes need their own schemas. CSIG carries compact bottleneck summaries, not an interchangeable full-hop trace. Budget added packet bytes, sampling rate and collector traffic; repeat the load test with capture disabled to expose observer effects. Hand the paired traces and acceptance bounds to Chapter 13.7.
Ultra Ethernet’s June 2025 public specification gave the merchant Ethernet ecosystem a common way to describe packet-level multipath, endpoint delivery/reordering, congestion signaling and supported loss recovery. That is a concrete alternative to qualifying only an integrated Spectrum-X or InfiniBand stack. The maintained UE release below separates network signals from receiver credits; choose a version-matched endpoint, switch profile and delivery mode, then test the same operations, traffic and faults. The multi-vendor bet pays when a second implementation completes that contract; publishing a mechanism does not measure job goodput.
MRC adds a distinct open transport path: its May 2026 production report describes an OCP RoCEv2 extension with packet spray, out-of-order placement and selective retransmission used by OpenAI and Microsoft. That is a named deployment record, with MRC 1.0’s Write-oriented operation boundary detailed in Chapter 8.4, not a substitute for UET’s broader contract. Spray, trimming and offload trade endpoint state, switch support, integration effort and recovery behavior. Keep the standards trigger in Chapter 16.2 and requalify when the version or supported operation changes.
Scope & caveats
Official specification history lists 1.0.3 as current; no public 1.1 established. Initial 1.0 release retained separately.
Choose the control that bounds the measured bottleneck and survives the lost path: movement for spare interior capacity, admission for a destination constraint, receiver repair for exhausted resources, and offload for eligible critical work. Choosing a mechanism from a port-utilization graph alone can move the queue while leaving the same GPUs waiting.
Cite this chapter
Fehn, J. (2026). Congestion Control, Load Balancing & In-Network Compute (Chapter 8.6). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-6-congestion-control-load-balancing-and-in-network-compute (accessed 2026-09-29).
@misc{aidc-8-6,
author = {Fehn, Jacob},
title = {Congestion Control, Load Balancing & In-Network Compute (Chapter 8.6)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-6-congestion-control-load-balancing-and-in-network-compute},
note = {Accessed 2026-09-29}
}