Chapter 8.4
In this chapter · 7 sections
Scale-Out Fabric: Protocols, Standards & Transport
A scale-out protocol is a 3–5 year bet on who supplies your switches, how much link rate survives as goodput, and whether you can leave the vendor whose firmware you depend on.
What you'll decide here
- Which back-end transport — InfiniBand, tuned RoCEv2, NVIDIA Spectrum-X, Ultra Ethernet/UET, or the OCP MRC extension to RoCEv2 — you standardize on, knowing the choice sets your switch/NIC suppliers and operational skill base. Reject missing Reads, Writes, Sends, atomics, ordering or completion semantics before comparing throughput, integration cost and recovery.
- Whether you accept a PFC + ECN/DCQCN lossless-Ethernet design and its head-of-line-blocking and deadlock failure modes, or select a supported lossy multipath transport that accepts out-of-order packets at the NIC — qualify loss recovery, duplicate handling and receiver resources in either mode.
- How much you are willing to pay — in switch premium, NIC dependence and integration effort — to close the gap between raw link rate and delivered all-reduce completion. Compare identical operations, placement and failures, including the retry cost of a classic in-order RoCE receiver where that is the selected implementation.
- Whether the fabric is single-tenant (a captive training supercomputer) or must enforce hard multi-tenant isolation (PKeys, VXLAN/EVPN, DPU-enforced VPC) — a decision that constrains protocol choice and the security boundary at once.
- Which parts of the stack you are willing to leave proprietary-and-fast today and re-decide as an open standard (UEC) matures — i.e. where you buy InfiniBand or Spectrum-X now and keep an Ethernet exit open.
A scale-out fabric has one job: carry the collective communication of a distributed training or inference job — all-reduce, all-gather, reduce-scatter, all-to-all — across the back-end network that stitches scale-up domains together, without becoming the bottleneck that idles the most expensive silicon in the building. The accelerators are bought and the power is contracted; the only variable left is how much of every step goes to computing versus waiting on the network. That fraction depends on the transport implementation, the paths it can use, and the endpoints’ ability to complete the required operations.
The decision presents as five options — InfiniBand, RoCEv2 in its selected lossless or lossy mode, NVIDIA Spectrum-X, Ultra Ethernet (UEC), and, since May 2026, the OpenAI-led MRC — but it is really one question asked twice. First: do you run a purpose-built lossless transport (InfiniBand) or Ethernet? Then, if Ethernet: does a RoCEv2 NIC's documented ordering, retry and PFC behavior meet the job's deadlines, or does a qualified multipath implementation — Spectrum-X, UET or MRC, each taken at its current published specification — justify its integration cost by spreading traffic across paths while preserving every required operation and completion rule? Each answer cascades into a switch-vendor list, an operational skill base, a congestion-control parameter space, and — the part nobody prices at scoping time — a multi-year lock-in to whoever owns the firmware your goodput silently depends on.
The protocol war: five answers to one question
For most of a decade the scale-out question had two default answers: InfiniBand for frontier training, and RoCEv2 on Ethernet for everyone cost-sensitive. That binary has shattered. Two forces broke it: the hyperscalers' refusal to single-source their largest capex line on one vendor's proprietary fabric, and the arrival of AI-tuned Ethernet that claims InfiniBand-class effective throughput on a merchant-silicon supply chain. The result is a genuine five-way fork, and the right answer now depends on cluster scale, tenancy model, in-house operational depth, and how much vendor lock-in you will tolerate against how much goodput you are willing to leave on the table.
InfiniBand (NVIDIA Quantum) is the incumbent for tightly-coupled training. It is lossless by construction — credit-based flow control means a sender never transmits a packet the receiver has no buffer for, so the fabric does not drop frames under congestion the way Ethernet does. It carries native RDMA, adaptive routing, and in-network reduction via SHARP (collective offload). The cost is a single-vendor supply chain (NVIDIA end-to-end: switch, NIC, cable, subnet manager), a separate operational discipline most Ethernet teams do not have, and a price premium. → in-network compute and SHARP are engineered in Chapter 8.6; the switch and NIC silicon in Chapter 8.3.
RoCEv2 (RDMA over Converged Ethernet v2) puts RDMA semantics on a routed UDP/IP Ethernet underlay. It promises the commodity economics and multi-vendor supply of Ethernet with RDMA's CPU-bypass data path. The catch is that classic RoCE inherited the go-back-N assumption of an ordered, lossless link: in the classic go-back-N implementation a gap can trigger retransmission from that point. PFC-based losslessness is one supported mitigation; a documented lossy RoCE mode with appropriate endpoint recovery is another. That engineering is the hard part, and the failure modes (head-of-line blocking, PFC deadlock, victim flows) recur throughout this chapter. Meta runs production RoCE at 24k-to-100k-GPU scale, which proves it works — but their published account makes clear how much fabric-engineering investment that took.
NVIDIA Spectrum-X is Ethernet that behaves like a purpose-built AI fabric: a Spectrum switch plus a BlueField/ConnectX SuperNIC — a version-matched Spectrum switch and ConnectX/BlueField endpoint combination, with Spectrum-4/5 deployment records kept separate from Spectrum-6 and ConnectX-9 component production and NVIDIA’s May 31, 2026 forecast of fall Vera Rubin system shipments (Chapter 8.3) — doing adaptive routing, per-packet spraying, NIC-side reordering, and programmatic congestion control to use multiple paths without the classic in-order penalty, subject to the supported endpoint contract. NVIDIA markets ~95% effective throughput versus the ~60% an untuned vanilla-RoCE fabric can collapse to under all-to-all collisions, and xAI's Colossus is the vendor-reported deployment behind that number, measured on the Spectrum-4 generation — a case study, not a cross-vendor benchmark. The cost: it is Ethernet on the wire but still effectively a NVIDIA end-to-end story — the switch and the SuperNIC come from one vendor, so the lock-in is real even though the protocol is nominally open.
Ultra Ethernet (UEC) is the industry's open answer — a full transport stack (UET) specified by the Ultra Ethernet Consortium, whose initial public 1.0 specification was released on 2025-06-11. It is designed from scratch for AI/HPC: packet spray across multipath, out-of-order delivery decoupled from message ordering with reassembly in the NIC, modern congestion control (UCCM), packet trimming for fast loss signaling, and native RDMA. The promise is a common transport contract for a multi-vendor supply chain, with AMD, Broadcom, Cisco, Arista, HPE, Intel, Meta and Microsoft among the UEC participants. The procurement cost is qualifying the maintained version across actual endpoints and switches: a compliance declaration does not prove that two suppliers complete the same operations or recover the same failure together. → standards trajectory consolidated in Chapter 16.2.
MRC (Multipath Reliable Connection) is the newest branch and the pragmatist's open path: an OCP specification (Release 1.0, dated March 2026) publicly released in May 2026 by OpenAI with AMD, Broadcom, Intel, Microsoft, and NVIDIA — and, unusually for a day-one standard, already production-proven: it ran on OpenAI's and Microsoft's largest training clusters before publication, including a 50k-GPU production pretraining job at OpenAI and a documented 75k-GPU pretraining job (consortium paper, 2026). Where UET replaces the RoCE stack outright, MRC extends RoCEv2 with the minimum machinery AI training needs: the data plane narrows to RDMA Write/Write-Immediate, every packet carries an entropy value so a single connection sprays across many paths, or whole planes; reliability moves to selective acknowledgment, with packet trimming as an optional fast loss signal where switches support it, and the fabric runs lossy with PFC off. Optionally, SRv6 source routing replaces dynamic path selection outright: hosts stamp the path onto the packet and the switches forward from static tables. The bet is UET's core performance ideas at a fraction of its scope — simpler to implement on existing RoCE-class NICs and merchant switch silicon, but narrower: a training-first transport, not a general HPC/storage stack. Vendor support landed fast on both ends of the wire (ConnectX-8, AMD Pollara/Vulcano, and Broadcom Thor Ultra among NICs; Tomahawk-class, Spectrum-X, and P4-programmable silicon among switches). Its multiplanar consequence — eight 100G planes from one 800G NIC reaching a 131,072-GPU two-tier fabric — is engineered in Chapter 8.5.
| Contract line | InfiniBand RC | RoCEv2 RC | Spectrum-X | UET / Ultra Ethernet | MRC 1.0 |
|---|---|---|---|---|---|
| Operations | Send/receive, Read/Write, atomics by device | RC verbs by NIC/driver; GPU registration separately | Supported RoCE/extension operation set by release | Profile/version and endpoint feature declarations | Write and Write-with-Immediate only |
| Ordering / completion | QP and verb rules; explicit remote-consumer synchronization | QP/verb rules; registration and fencing matter | Enhanced placement must preserve the chosen API contract | Select delivery mode and profile; test completion visibility | OOO placement; transport completion and WriteIMM resources |
| Loss / duplicates | Link credits prevent congestion drops; errors still require recovery | Classic go-back-N or documented NIC extension; test loss/reorder | Compatible endpoint retries/reorder; no generic latency promise | Declared loss recovery, retry and duplicate handling | Selective retransmission; no Read/Send/Atomic; no RNR-NAK |
| Congestion / path | Credit, routing and virtual-lane configuration | PFC/ECN or qualified lossy mode; path granularity by endpoint | Integrated switch/NIC features in named firmware | Versioned endpoint/network signals and profile requirements | Lossy multipath; trimming optional; controller config required |
| Tenant boundary | PKeys plus memory permissions; confidentiality separate | Reachability plus keys/IOMMU; confidentiality separate | Same separate protection and performance tests | Profile security/features plus implementation tests | Memory keys plus independent reachability/performance policy |
| Supply / support | Selected NIC, switch, manager and library combination | Multiple suppliers only where the combinations interoperate | Integrated feature support can constrain endpoint substitution | Public spec; obtain matched declarations and interoperation | Open OCP spec; operation scope restricts substitution |
| Evidence / date | Selected versioned device and verbs manual | Selected versioned NIC/NOS documentation | Named vendor system release and workload result | 1.0: 2025-06-11; maintained 1.0.3: 2026-07-16 | OCP document 2026-03-21; public release May 2026 |
Transport semantics: loss, ordering, completion and failure
Underneath the five product names sit two orthogonal semantic axes that actually determine behavior: does the fabric drop packets under congestion (lossy) or refuse to (lossless)? and must packets arrive in order, or can the endpoint reassemble an out-of-order stream? Every transport is a point in that 2x2, and the modern AI-Ethernet entrants exist precisely because the historically dominant corner — lossless-and-ordered — turned out to be a trap at scale.
Lossless means the fabric exerts back-pressure to prevent buffer overflow rather than dropping frames. InfiniBand does this natively with credits. Ethernet retrofits it with PFC (IEEE 802.1Qbb): when an ingress buffer fills, the switch sends a PAUSE upstream for that traffic class, stopping the sender. The problem is that PFC is a blunt, per-class, per-link hammer — it pauses all flows in the class on that link, not just the one causing congestion. That is head-of-line blocking: an innocent flow sharing the link is paused for a congested neighbor it has nothing to do with. Worse, PAUSE propagates hop-by-hop upstream, and in a multi-tier Clos with a cyclic buffer-dependency it can deadlock the fabric entirely — a class of failure that needs watchdogs to detect and break. → the PFC/ECN/DCQCN parameter space and its pathologies are the subject of Chapter 8.6.
Ordered delivery is the second trap. Classic RoCE assumes a packet stream arrives in sequence; an out-of-order arrival is treated as a loss, triggering go-back-N retransmission — the receiver discards everything after the gap and the sender retransmits from the lost packet forward. This is the RoCE in-order penalty, and it is why classic RoCE cannot freely spray a flow across multiple equal-cost paths: if two paths have different latency, the resulting reordering looks like loss and triggers cascading retransmits. So vanilla RoCE pins each flow to a single path (ECMP hashing on the flow tuple), which means a single hot link can bottleneck a flow while parallel links sit idle — exactly the all-to-all collision pattern that collapses untuned RoCE goodput.
Spectrum-X, UEC, and MRC all attack the same root cause: they decouple wire order from message order. The NIC sprays packets of a single flow across many paths, accepts them out of order, and places accepted packets and tracks message completion according to the operation’s ordering contract. A delay, a bit of congestion, or a single dropped packet on one path no longer poisons the whole flow — selective retransmission replaces go-back-N, and selective retry can reduce retransmission to the missing packet rather than repeating the later window, while feedback, retry and outstanding dependencies still determine the completion delay. UEC adds packet trimming: instead of dropping a congested packet outright, the switch truncates it to its header and forwards the stub, so the receiver learns of the loss immediately and signals a fast, surgical retransmit. This decoupling of wire order from message order is why the protocol war is really a war over who owns the NIC that does the reordering.
First compare required operations with implemented operations. A and B are eligible for qualification because their assumed device declarations include Read, Write and Write-with-Immediate. C is excluded before throughput comparison: the MRC 1.0 specification dated 2026-03-21, §5.1 and §6.2.2, excludes Reads, Sends and atomics. Changing a NIC’s rate cannot supply a missing verb. The flip is the workload’s Read requirement: if every remote fetch is safely converted to an owner-pushed Write and its tested notification/ownership protocol, C becomes eligible too. That conversion costs application work and receiver-resource management; it is not a transparent transport swap.
Then run the completion trace. In the successful transfer, register and authorize the destination, post the receiver resource, issue the Write-with-Immediate, accept all three packets, and release the consumer only after the required receiver completion and memory synchronization. A sender completion alone does not prove that an application on the remote accelerator has consumed the bytes. If the middle packet is lost, the selected reliable implementation retries under its documented rules; incomplete data must not become a successful consumer event. A duplicated packet must not produce a second logical notification or repeat a non-idempotent application action. After retry exhaustion, require the documented completion error, stop dependent consumers, invalidate uncertain state and restart the application protocol from an agreed boundary.
The RDMA programming manual supplies the RC verbs context. No result is measured here: selection is HOLD between eligible candidates until the same operation, visibility, loss, duplicate and failed-completion tests pass, followed by Chapter 8.1’s workload comparison. Pick the passing implementation whose support and substitution cost the operator can carry; Chapter 13.7 records installed acceptance.
The RoCE in-order penalty, and what the published numbers do not settle
The headline number that drives the entire AI-Ethernet movement is the gap between tuned and untuned RoCE effective throughput. ECMP flow-pinning can collide elephant flows on one path while others idle; congestion, drops and a classic go-back-N implementation can then amplify completion time. Quantify the penalty on that exact endpoint, algorithm and traffic matrix. NVIDIA reports Spectrum-X at Colossus at roughly 95% data throughput, contrasting a 60% standard-Ethernet case. That vendor-reported deployment result does not establish a generic tuned-RoCE value; validate the named workload, topology, transport, endpoints, and measurement boundary. The gap decides real money: at 100,000-GPU scale the network is the second-largest line after the accelerators themselves, and NVIDIA's 2024-10-28 figures of 60% and 95% data throughput set that vendor's fabric against an unnamed standard-Ethernet baseline, not against a named competitor. What that difference costs a buyer, on the same job, is the accelerator time lost when exposed transfers miss the completion target.
The consequence for the decision: RoCE's cost advantage is real only if you have the team to capture it. The merchant-silicon switch is cheaper than InfiniBand, but the tuning labor, the validation rigs, and the on-call fabric expertise are not. Spectrum-X sells that tuning as an integrated product, while UEC standardizes related mechanisms for multi-vendor implementations; neither removes the need to validate delivered goodput on the target fabric, so the operator can reuse a qualified congestion design, while still owning endpoint settings, fault diagnosis and release qualification. You are, in effect, choosing whether to buy the goodput as a capability (Spectrum-X/UEC) or build it as a competency (hardened RoCE).
Scope & caveats
Historical initial public UEC 1.0 release. The maintained specification revision is recorded separately; publication alone does not qualify an implementation.
Scope & caveats
Historical 8-byte host-memory MPI ping-pong mean, not p99, GPU-memory latency or a RoCE comparison. Daytona_X / EPYC Rome / ConnectX-6 HDR; OSU 5.6.2, HPC-X 2.7.0, OFED 5.0.2; local core 80, 10,000 timed/warm-up iterations.
Scope & caveats
NVIDIA-reported Colossus deployment result under that workload and configuration, not a universal Ethernet, Spectrum-X, or scheduled/VOQ-fabric figure; Colossus is an adaptive-routing/telemetry RoCE fabric, not the scheduled-fabric category. Meta reports tuning RoCE and InfiniBand GenAI clusters to equivalent performance — no common-workload test crowns either transport.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Scope & caveats
Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.
Scope & caveats
named reference designs and deployments with different traffic, topology, placement and service objectives
Derive oversubscription from the measured traffic matrix, collective/request mix, topology, failure headroom and SLO; validate it on the target fabric.
Scope & caveats
The OCP spec document (Release 1.0) is dated March 21, 2026; OpenAI's public release was May 5, 2026, with partner announcements May 6. Production use is consortium-reported: the companion paper (arXiv:2605.04333) documents MRC in production at OpenAI and Microsoft, including a 50k-GPU OpenAI production pretraining job and a 75k-GPU pretraining job (organization unattributed).
Standards trajectory: where the protocols are heading
The strategic shape of Ethernet in 2026 is a choice between integrated platforms and open implementation paths. InfiniBand couples its transport with supported SHARP offload for tightly coupled runs, while Spectrum-X gives NVIDIA an Ethernet platform and UEC gives the merchant ecosystem a specification to implement; none of those labels establishes a throughput winner for an untested job. The initial UEC 1.0 release on 2025-06-11 is the historical inflection: it describes a coordinated stack across physical, link, transport, software, storage and management, giving NIC and switch suppliers common interfaces against which to build. The sourcing bet pays off when a second implementation meets the same operation and failure contract, at which point the buyer can compare its integration cost with the incumbent's price and support obligation. Two developments make that choice concrete. The maintained specification is 1.0.3, released 2026-07-16, with no public 1.1 established by the release history; Broadcom's 800G Thor Ultra announcement of 2025-10-14 claims UEC feature compliance, while AMD's August 2025 Pollara 400 presentation describes a distinct 400G implementation, so neither the NIC rates nor compliance labels are interchangeable. MRC supplies another open path: its May 2026 production report describes a RoCEv2 extension with packet spray, out-of-order placement and selective retransmission in the reported system. The two open transports therefore address different operation and integration boundaries: UET has the broader contract, while MRC 1.0 excludes Reads, Sends and atomics. A workload that needs one of those operations pays for application changes before MRC becomes eligible; a buyer choosing UET pays for version-matched qualification because the UEC compliance process uses self-attestation, not a guarantee that arbitrary devices interoperate.
For the decision-maker, this trajectory argues for explicit optionality management. If you must deploy at frontier scale on a deadline that favors an already supported integrated fabric, buy that combination after workload acceptance and capture the completion-time gain it actually demonstrates — but build the cluster’s physical layer (cabling, optics, structured plant) for the channel, lane, connector and module requirements of each credible Ethernet/InfiniBand replacement, so the transport can be re-decided at refresh after qualification. The physical investment is the irreversible part; the protocol running over it is comparatively reversible if you did not hard-wire your topology to one vendor's switch radix. → the full subsystem roadmap, including UEC milestones and the 800G→1.6T→3.2T optics ladder, is consolidated in Chapter 16.2; the physical-layer choices that preserve or destroy that optionality are in Chapter 8.9 and Chapter 8.10.
UE profile, coexistence and qualification
Freeze the UET specification version, software API and delivery mode with the NIC/switch/firmware set. The 1.0.3 specification uses AI Base, AI Full and HPC profiles; the older compliance readme uses AI Extended. Match the artifacts by edition and feature rows rather than silently equating profile names. Mark Reads, Writes, Sends, atomics, matching, ordering, retry limits, security and receiver-resource limits as required, supported or absent for the actual endpoint. An optional mechanism on a data sheet is not enabled merely because both devices use Ethernet.
For coexistence with RoCE or ordinary traffic, document classification from endpoint marking through switch queue and pool, MTU/FEC, routing entropy, trimming behavior where enabled, congestion signals and which traffic can be paused. Test load and failure with the coexisting traffic present, including receiver-resource exhaustion and a peer that lacks an optional feature. Separate reachability from interoperation: packets reaching an endpoint do not prove that its transport recognizes the operation or completes it correctly. Migration requires the old and new application paths to have explicit owners; there is no assumed wire compatibility between RoCE, UET and MRC. Chapter 8.6 owns congestion/reduction tuning, and Chapter 13.7 owns the installed evidence record.
Scope & caveats
Official specification history lists 1.0.3 as current; no public 1.1 established. Initial 1.0 release retained separately.
Deep dive: why packet spray + NIC reorder wins (and what it costs)
Strip away the brand names and the entire AI-Ethernet revolution reduces to one mechanism: spray a single flow's packets across every available path, accept them out of order, and reassemble in the NIC. Understanding why this is so powerful — and what it costs — is understanding the chapter.
Why it wins. AI collectives generate a small number of enormous "elephant" flows (a single GPU's all-reduce contribution can be gigabytes). Classic ECMP hashes each flow to one path, so two elephants can collide on one link while seven parallel links sit empty — and because RoCE is in-order, you cannot simply split the elephant across the seven free links without triggering go-back-N. Packet spray breaks the flow into per-packet (or per-"entropy") units load-balanced across all paths, so an elephant uses the full bisection bandwidth and a flow can escape a congested link through alternate paths instead of remaining pinned there, although one missing packet can still delay its dependent completion. This mechanism helps explain NVIDIA's reported ~95% Colossus result, but the number remains workload- and configuration-specific and is not a generic sprayed-fabric guarantee.
What it costs. The reassembly is not free. The NIC must place each packet straight into destination memory from the address and key it carries, track loss and duplicates, and hold per-message completion state until the message is whole — placement out of order rather than a serial reorder buffer, but state and logic at line rate all the same, which is precisely why these are SuperNICs (ConnectX/BlueField for Spectrum-X) or UET-compliant NICs, not commodity Ethernet adapters. That is the lock-in vector: the protocol is open-ish, but the NIC that makes it fast is sophisticated, and the switch and NIC must agree on the spraying/trimming scheme. So "open Ethernet" still means "a NIC and switch that implement the same advanced transport," which in 2026 is a short list of suppliers. UEC's value is that it standardizes the contract between switch and NIC so that, eventually, a Broadcom switch and an AMD NIC interoperate — turning today's vendor-pair lock-in into tomorrow's mix-and-match. Until that interop is proven in production, each untested substitution carries integration work that a previously qualified combination has already completed.
MRC reaches the same destination through a narrower door: instead of a new transport stack, a RoCEv2 connection whose packets each carry an entropy value for per-packet spraying, with selective acknowledgment and reassembly at the receiver — spray without leaving the RoCE ecosystem, at the cost of supporting only the write-style operations AI training uses (→ the MRC branch above).
Multi-tenant fabrics: isolation as a fabric concern
Everything above assumes a single-tenant supercomputer. The moment the fabric serves multiple customers — any GPU neocloud, any cloud GPU service — the transport choice collides with an isolation requirement, and the two cannot be decided separately. A multi-tenant back-end fabric must stop one tenant from reaching or reading another tenant’s traffic and constrain the congestion or attacks it can impose, enforcing reachability, memory authority, confidentiality and promised performance isolation as separate controls, and it must do so on a high-bandwidth fabric whose required completion and loss-recovery behavior constrains how buffering and policing can isolate tenants.
The mechanisms differ by transport. InfiniBand isolates with partition keys (PKeys) enforced by the subnet manager — a tenant's nodes share a PKey and cannot communicate across partition boundaries. Ethernet fabrics overlay tenancy with VXLAN/EVPN (network virtualization that gives each tenant an isolated L2/L3 segment over a shared underlay). One enforcement location outside a tenant-controlled host is the DPU: a BlueField/Pensando-class DPU at the host edge terminates the tenant's VPC, enforces microsegmentation and encryption, and rate-limits per-tenant — so isolation is enforced in hardware at the boundary of every server rather than trusted to the fabric core. DPU-enforced VPC isolation and VLAN/VXLAN segmentation solve different parts of tenancy: the isolation test must prove which policy, DMA and memory controls survive tenant-root compromise. An overlay identifier or hardware device alone does not establish that boundary.
The consequence for protocol choice: multi-tenancy can favor Ethernet + DPU when the DPU enforces the required VPC, DMA and memory boundary outside tenant control. InfiniBand PKeys give partitioning but the DPU-enforced VPC model — full network virtualization, per-tenant encryption, line-rate microsegmentation — maps more naturally onto an Ethernet fabric with programmable DPUs at the edge. A neocloud that must rent isolated slices to mutually-distrusting tenants on shared GPUs is therefore pulled toward the Spectrum-X / UEC + DPU side of the fork, where the isolation primitives and the security tooling are richer — even before goodput enters the conversation. This is where the fabric decision and the security decision merge: the same DPU that enforces the tenant boundary is the zero-trust enforcement point. → the tenant security boundary is engineered in Chapter 11.6, and zero-trust microsegmentation in Chapter 11.7; the DPU silicon itself in Chapter 8.3.
| Mechanism | Fabric | What it isolates | Strength | Cost / limitation |
|---|---|---|---|---|
| Partition keys (PKeys) | InfiniBand | Membership — who can talk to whom | Strong partitioning, SM-enforced | Coarse; no per-tenant encryption or VPC semantics |
| VLAN | Ethernet | L2 broadcast domain | L2 segmentation; requires correct configuration and policy | Does not scale to cloud tenancy; trivial to misconfigure |
| VXLAN / EVPN | Ethernet | Virtualized L2/L3 per tenant over shared underlay | Scalable overlay isolation | Underlay still shared; congestion isolation imperfect |
| DPU-enforced VPC | Ethernet (+ DPU) | Full per-tenant VPC, microseg, encryption, rate-limit | Hardware enforcement outside tenant OS when correctly owned | Requires a DPU per host; cost and integration |
Deep dive: the operational-skill tax nobody scopes
The line item that wrecks fabric decisions is not on any vendor quote: the operational-skill tax. Each transport demands a different competency, and the cluster's reliability is hostage to whether you have it.
InfiniBand needs a subnet-manager discipline: managing the SM, partitioning, routing tables, and ibdiagnet-driven BER validation — a skill set that lives mostly in HPC shops and inside NVIDIA's ecosystem, scarce in general cloud-ops teams. RoCE needs a congestion-control research competency: someone who genuinely understands PFC thresholds, ECN marking points, DCQCN dynamics, and how they interact at your scale and topology — and who can debug a PFC deadlock from telemetry at 3 a.m. Spectrum-X trades much of that for vendor-managed behavior, but ties you to NVIDIA's switch-and-SuperNIC support model and its firmware cadence. UEC, in its current maturity, demands the most integration skill — making a multi-vendor NIC/switch combination interoperate and perform — which is exactly why its early adopters are hyperscalers with deep in-house networking teams, not enterprises.
The decision rule: pick the transport whose skill tax you can actually pay, not the one with the best spec sheet. A cost-optimal RoCE fabric operated by a team without congestion-control depth will under-deliver goodput so badly that the cheaper switch becomes the more expensive cluster. Conversely, a hyperscaler with a standing fabric-engineering function can extract RoCE/UEC's economics that an enterprise cannot. The fabric is only as good as the team that tunes it — and that team's cost belongs in the BOM. → fabric commissioning and the validation gates that catch a mis-tuned fabric before production are in Chapter 13.7; congestion telemetry and observability in Chapter 10.6 and Chapter 14.2.
Putting it together: how to choose
The choice resolves along four rules, in order. (1) Scale and coupling. Frontier, tightly-coupled training that lives and dies on all-reduce goodput, with a team fluent in HPC fabrics, still defaults to InfiniBand — the SHARP offload and in-order losslessness remain the highest-goodput known quantity, provided the supported switch/NIC/library combination is the one you buy. (2) Supply-chain strategy. If single-sourcing your largest network capex on one vendor is strategically unacceptable, you are on the Ethernet side of the fork regardless of goodput — the only question is which Ethernet, and whether the second supplier’s implementation actually inter-operates. (3) Time vs maturity. If you need near-InfiniBand goodput on Ethernet today, Spectrum-X is the finished product and you accept the NVIDIA switch+SuperNIC pairing; if you can absorb integration risk to own a multi-vendor future, UEC is the bet; if the workload is training-shaped and you want open multipath on current RoCE-class silicon, MRC is the narrower open bet — already carrying frontier production traffic, though a required Read verb excludes MRC 1.0. (4) Tenancy. Hard multi-tenant isolation pulls toward Ethernet + DPU, where VPC virtualization and microsegmentation are native and the memory, confidentiality and performance controls that must survive tenant compromise live. Whichever rule selects, the tested combination must meet the same healthy and degraded workload deadline, recover from loss and duplicates without releasing an incomplete result, and name who owns the pager and firmware lifecycle.
Two anti-patterns recur. The first is buying link rate instead of goodput — specifying a faster port and getting a slower cluster because the transport collapses under collisions; the fix is to benchmark delivered all-reduce bandwidth, not read the port label. The second is adopting RoCE without the team to tune it — capturing Ethernet's switch discount while forfeiting its goodput, so the “cheaper” fabric misses its communication deadline and idles the accelerators it was supposed to feed. Both come from deciding on the spec sheet rather than on delivered goodput and total operational cost — so price the fabric, validate it, and staff it against the goodput it has to deliver. → the topology and oversubscription decisions that sit on top of this transport choice are in Chapter 8.5; the traffic characterization that motivates all of it in Chapter 8.1; the scale-up domain this fabric stitches together in Chapter 8.2.
Cite this chapter
Fehn, J. (2026). Scale-Out Fabric: Protocols, Standards & Transport (Chapter 8.4). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-4-scale-out-fabric-protocols-standards-and-transport (accessed 2026-09-29).
@misc{aidc-8-4,
author = {Fehn, Jacob},
title = {Scale-Out Fabric: Protocols, Standards & Transport (Chapter 8.4)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-4-scale-out-fabric-protocols-standards-and-transport},
note = {Accessed 2026-09-29}
}