The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 8.3

In this chapter · 9 sections
Term help

Network Silicon: Switch ASICs, NICs & DPUs

The switch ASIC, NIC and DPU set the ceiling on every fabric above them, so their SerDes generation, buffer architecture and offload engines must close with the topology before the purchase is frozen.

GOODPUTPOWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. Which switch ASIC family you standardize on — merchant high-radix Broadcom Tomahawk, merchant deep-buffer Broadcom Jericho, Cisco Silicon One, or NVIDIA Quantum InfiniBand / Spectrum Ethernet — because the supported ASIC + NIC + NOS combination sets radix, usable buffers, queues, telemetry, host attachment, populated-switch heat and vendor lock-in for the cluster.
  2. Whether your back-end NIC is a plain RoCE/IB NIC or a SuperNIC with full transport offload — and whether the host even needs a DPU, or whether a NIC suffices, because the DPU is a per-server tax you justify with storage, security, and multi-tenant isolation, not with raw bandwidth.
  3. Shallow shared-buffer vs deep-buffer VOQ at each tier: shallow when usable queues absorb the admitted burst until PFC/ECN or admission acts, deep when the larger bounded burst justifies its power and queueing cost. Crossing buildings or oversubscribing changes the feedback budget, but neither buffer architecture rescues sustained overload.
  4. Which functions you push off the host CPU and onto the DPU (storage initiator, encryption, the VPC overlay, the security policy plane) versus leaving on x86 — every offloaded function frees host cores for the workload but adds a second control plane to operate and patch.
  5. Whether the chosen SerDes and port modes carry you through a GPU refresh, because lane rate helps set port count, reach and the copper-vs-optics break. 224G-class SerDes and 200G/lane Ethernet service describe the same interface generation at different boundaries; a higher-rate roadmap becomes useful only with supported MAC modes, FEC and a qualified channel.
Compare Broadcom, Cisco and NVIDIA ASIC systems and NICs on the same MAC, electrical-lane/SerDes, PHY/FEC, cage/breakout, queue/NOS and endpoint-operation axes. The bound chip and NIC records establish separate product interfaces; they do not establish interoperability. Buy only a combination whose port ledger and qualified software realize the required topology, or accept the integration and replacement cost of changing the system.

The fabric chapters that bracket this one — scale-up in Chapter 8.2, scale-out protocols in Chapter 8.4, topology in Chapter 8.5 — all assume a set of silicon building blocks and reason about how to wire them together. This chapter is about the blocks themselves: the three classes of programmable silicon the entire AI network is assembled from. The switch ASIC moves packets between ports. The NIC (and its beefed-up cousin, the SuperNIC) connects a server's accelerators to the wire and runs the RDMA transport. The DPU/IPU is a NIC with a CPU complex bolted on, sitting in the data path to offload storage, security, and virtualization from the host. Select the silicon well and the fabric design that follows is a series of well-posed wiring problems; under-spec the SerDes, mismatch the usable buffer allocation to the admitted burst, or buy a DPU you have no offload for, and you have baked a ceiling into the cluster that no topology cleverness can lift.

These decisions are unusually consequential because they come early and stay sticky. The switch ASIC family fixes your radix and your lock-in posture before a single cable is run. The SerDes generation on that ASIC fixes, for the next several years, how many ports you get, how far copper reaches, and when optics become mandatory. And the DPU decision — whether to deploy one at all, and what to run on it — is a per-server line item across thousands of servers that pays back only if you actually move work onto it.

The gating spec: SerDes generation, not aggregate Tbps

Marketing leads with “102.4 Tbps switch” or “800G NIC,” but the gating specification is the complete port contract: MAC service rate, electrical lanes, coded signaling, FEC, cage arrangement and supported breakout. The lane still governs how many interfaces fit and how far the electrical route reaches before optics add power and cost. A vendor’s 224 Gb/s-class SerDes label and 200 Gb/s-per-lane Ethernet service are not consecutive generations. Multiplying the electrical ceiling by lane count does not produce useful Ethernet throughput. Cisco’s G300 data sheet dated 2026-02-10 lists those electrical and MAC specifications separately. IEEE P802.3dj remains a draft project; specify the interface document and FEC mode rather than declaring a product port to be a ratified application.

Why does the lane matter alongside the aggregate? Broadcom’s Tomahawk 6 launch illustrates the slicing: its 2025 announced 102.4 Tb/s MAC capacity supports 64 × 1.6T, 128 × 800G or 256 × 400G logical-port modes, while a purchased box implements only some of those arrangements. An optical module can also gearbox between different electrical and optical lane counts. Count ASIC-to-module lanes, optical transmit lanes and active fibers separately; none is the number of faceplate cages. Read supported modes from the switch and module release together, then confirm the host/NIC path reaches them. The copper cliff remains a route constraint, not a universal meter count: lower-rate electrical lanes can improve a qualified copper route while consuming more lanes per port, so “newer SerDes” is not a complete reach or power decision. The physical-layer consequences belong in Chapter 8.9; the transport operations implemented at the endpoints belong in Chapter 8.4.

Switch ASIC families: the supported system behind the silicon fork

Switch silicon differs on buffer organization, forwarding features and support boundary as well as aggregate rate. Broadcom Tomahawk and Jericho, Cisco Silicon One G300, and NVIDIA Spectrum/Quantum all belong in the same comparison. Buffering philosophy decides where bursts wait, while the business model decides who controls support and replacements. Separate the ASIC supplier from the system supplier and the NOS: a chip’s SDK feature does not prove that a purchased software release exposes it. Freeze one supported combination of SKU, firmware, driver, NOS, transceiver and collective library, then ask which row prevents that combination from carrying the topology. A missing queue-allocation or telemetry line is a contract gap even when every headline bandwidth number passes.

Broadcom Tomahawk is the merchant high-radix line: it favors on-chip shared buffering and endpoint-controlled traffic over storing prolonged overload in external memory. Broadcom’s 2026-03-12 release says the 102.4 Tb/s Tomahawk 6 family is shipping in production volume. That record follows the initial launch; it does not prove volume availability for every package or optical variant. PFC/ECN and endpoint admission still have to limit bursts before the shared queues fill. Demand the actual port map and the selected NOS’s queue, load-balancing and counter support before turning chip radix into leaf capacity.

Broadcom Jericho is the merchant deep-buffer, VOQ line: routing-class silicon that spends external memory on bursts too large for an on-chip queue budget. Broadcom’s August 2025 Jericho4 announcement advertises a 51.2 Tb/s family aggregate and HBM-backed packet buffering; those resources buy time for a bounded arrival/service mismatch, not more sustained capacity at the destination link. Compare packet-routing mode with scheduled-fabric mode, including fabric elements, admission protocol, power and backlog tail. Longer routes enlarge the feedback budget, but even a deep buffer eventually fills under sustained overload. Chapter 8.8 admits cross-site traffic against the protected path.

NVIDIA Quantum and Spectrum provide InfiniBand and Ethernet system paths. Quantum-X800’s 800G/port InfiniBand with SHARPv4 in-network reduction — 14.4 TFLOPS of in-network compute, nine times the prior generation (NVIDIA, March 2024; mechanics in Chapter 8.6) — and Spectrum-X paired with ConnectX/BlueField endpoints are concrete integrated candidates in the dated NVIDIA records. Their collective offload and endpoint/switch features have to survive the exact software combination being purchased. Cisco G300 supplies another current Ethernet candidate on the same axes. A controlled workload comparison can justify an integrated premium; NVIDIA’s October 2024 Colossus report of roughly 95% data throughput is one deployment result, not a transport winner. The economic choice is whether one support boundary saves enough integration and recovery effort to outweigh the loss of interchangeable components. Chapter 7.1 develops the merchant-versus-captive business model.

ASIC families → the supported system contract
Contract lineTomahawk 6Cisco G300Spectrum-6 classJericho classQuantum-X800
MAC capacity / modes102.4 Tb/s family; system must expose selected modes102.4 Tb/s; 64×1.6T in data sheet102.4 Tb/s ASIC in SN6000 manualNamed SKU and forwarding mode requiredNamed InfiniBand switch port configuration required
Electrical lanes512×200G or 1,024×100G options512×224G-class SerDes; separate from MAC rateUse selected SN6000 electrical port modeExact SKU lane map requiredExact switch/NIC link mode required
Cages / opticsSystem and package-specific; qualify each populated modeSystem cage map required; ASIC radix is not cage countSN6600-LD manual: 64 OSFP cages; mode-specific portsLine card and fabric-module map requiredPhysical cages and logical ports counted separately
Buffer / poolShared on-chip; obtain usable allocationFully shared on-die; allocation depends on policyManual: 160 MB per ASIC shared; verify usable allocationVOQ/external memory by SKU and system modeCredit-managed resources; obtain per-class limits
Queues / classesNOS class mapping, limits and pool policy requiredMultiple output queues; selected NOS limits requiredNOS class mapping and pool policy requiredVOQ destination/class map and credits requiredVirtual lanes, arbitration and management policy required
Path selection / reorderDocument enabled routing; endpoint reorder gates sprayECMP/weighted, flow/flowlet, spray; endpoint contractIntegrated routing depends on NIC/firmware combinationScheduled or packet mode; state whichSubnet-manager/routing configuration required
TelemetryCounters, microburst capture and timestamp origin requiredECN, PFC watchdog, flow/queue and event telemetry listedExpose port/queue/error and selected transport countersVOQ depth, grants, fabric health and endpoint progressPort/error, route and offload/fallback counters
NOS / supportSystem vendor and exact NOS releaseSDK/SAI and SONiC reference listed; system release requiredNamed NVIDIA/system software support matrixInterface/fabric elements and software as one contractSwitch, NIC, manager and collective release as one contract
Evidence date / statusBroadcom volume announcement 2026-03-12Cisco data sheet 2026-02-10; no job uplift inferredSN6000 manual, preliminary May 2026; supply by SKUObtain dated selected-system evidenceObtain dated selected-system evidence
Primary records: Broadcom production-volume announcement (2026-03-12), Cisco G300 data sheet (2026-02-10), and NVIDIA SN6000 hardware manual (May 2026 preliminary revision). Unspecified system limits are required contract inputs, not zeros or inferred features.
102.4 Tb/s; 64×1.6T; 512×224G-class SerDes
G300 data sheet, 2026-02-10
Qualify the system and software subset.
Scope & caveats

Chip specifications, not guaranteed system cages, usable buffers or application acceleration. Electrical SerDes class differs from MAC service rate.

102.4 Tb/s ASIC; 160 MB shared buffer
SN6000 manual, May 2026 preliminary
Separate buffer capacity from usable queue guarantees.
Scope & caveats

SN6000 manual revision 1.0, May 2026 (publication day unspecified; date normalized to month start). SN6600-LD has 64 OSFP cages with mode-specific logical ports; no shipment-date inference.

Shallow-shared vs deep-buffer VOQ: the buffering tradeoff

Tomahawk versus Jericho is the cleanest hardware fork in networking, and it has no free lunch. The question it turns on: where do you store a packet that arrives faster than its egress port can drain? Two silicon philosophies answer it differently.

Shallow shared-buffer ASICs (Tomahawk-class) keep a small, fast pool of on-chip SRAM shared across all ports. The bet is that with short reach, a non-blocking topology, and good endpoint congestion control (PFC/ECN/DCQCN, adaptive routing, packet spray — the machinery of Chapter 8.6), bursts are absorbed at the source and the switch never needs to hold much. The payoff is lowest latency, highest radix, and lowest power-per-bit. The risk: when a real incast burst exceeds the shallow buffer, you must either drop (lossy) or assert backpressure (PFC), and PFC at scale brings head-of-line blocking and deadlock risk. Shallow buffering only works if the congestion-control loop is fast enough to keep the buffer from filling.

Deep-buffer VOQ ASICs (Jericho-class) attach large off-chip memory — HBM on Jericho4 — and organize it as virtual output queues: a separate logical queue per egress destination, so a congested port cannot head-of-line-block traffic bound elsewhere. The payoff is the ability to absorb enormous bursts and to run lossless over long reach — a larger burst envelope for RoCE over long DCI routes when admission and feedback keep arriving bytes within it; depth alone cannot prevent drops. The cost is real: added latency (a packet may sit in deep buffer), higher power and die area (HBM is not free), and higher $/port. You do not want deep buffers on a short-reach leaf where they add latency you never needed; you do want them at the fabric edge that crosses buildings.

Close each line before choosing the leaf. The MAC total is 64×800 Gb/s=51.2 Tb/s per direction; downstream and upstream each carry 32×800=25.6 Tb/s. The cage map consumes 64×4=256 electrical lanes, while the optical map has 64×8=512 transmit lanes and 64×16=1,024 active fibers. Those counts are distinct. A 102.4 Tb/s ASIC can have enough aggregate capacity without the selected box exposing these cages or modes. A dual-400GbE-only adapter fails the single-800GbE attachment line even if its aggregate is 800 Gb/s.

The required usable queue allocation is 64×0.60=38.4 MB, inside the assumed 40 MB pool, but candidate A’s per-port cap of 0.512 MB fails. Reject that configuration before comparing throughput. The flip is a documented and tested per-port allocation of at least 0.60 MB while all 64 required queues are active and the other pool reservations still fit. At that threshold the arithmetic closes; purchase remains HOLD until the cage map, host payload, telemetry and concurrent-buffer test produce evidence. The G300 data sheet and SN6000 manual illustrate why advertised chip features and system allocations are separate. Chapter 13.7 accepts the installed combination.

NICs and SuperNICs: the RoCE/IB offload path

The NIC is where the network meets the accelerator, and in AI fabrics it does far more than push frames. The defining feature is RDMA — remote direct memory access — which lets a GPU on one node read or write registered GPU memory on another without host data copies when registration, PCIe/IOMMU topology and software support the path (with the CPUs retaining the required setup and control), the foundation that makes collective communication tolerable at scale. Two transports carry RDMA: native InfiniBand (NVIDIA ConnectX in IB mode) and RoCEv2 (RDMA over Converged Ethernet), which runs the same verbs over a routable Ethernet/UDP underlay. The NIC implements the transport in hardware; the quality of that implementation — how it handles congestion, retransmission, and packet reordering — is part of why one fabric advances the job while another wastes bisection capacity on retries or idle barrier wait; measure the selected job’s service rate and completion tail.

One supported back-end reference pattern is one NIC attachment per GPU: an 8-GPU server carries 8×400G or 8×800G back-end ports (3.2–6.4 Tb/s/node), plus a separate, smaller NIC for the front-end/storage/management plane. The term SuperNIC denotes the AI-optimized variant: full transport offload, hardware support for the adaptive routing / packet-spray and out-of-order reassembly that Ultra Ethernet and Spectrum-X require, and line-rate congestion handling. ConnectX-8 carries the current 800 Gb/s generation — with a port-structure subtlety that shapes fabric design: its single 800 Gb/s physical port runs Ethernet as 2×400GbE (a single 800GbE link is unsupported; the same port does carry 800 Gb/s XDR InfiniBand), which is why the B300-generation reference design wires each GPU into two independent 400GbE planes (NVIDIA, 2025–26; → the planes pattern in Chapter 8.5). The Rubin platform doubles per-GPU scale-out bandwidth to 1.6 Tb/s, delivered as two 800 Gb/s ConnectX-9 SuperNICs per GPU — eight per four-GPU compute tray — the 800 Gb/s C9180-class parts being the family's first with single-port 800GbE (a 400 Gb/s variant also exists; NVIDIA, 2026). The pair teaches the general rule: read a NIC's supported link speeds alongside its aggregate, because the two differ across the whole 2026 field. And the field is no longer single-vendor: AMD's Pensando Pollara 400 (announced October 2024 as the industry's first UEC-ready AI NIC, generally available April 2025) and 800G Vulcano, and Broadcom's Thor 2 and 800G Thor Ultra (full UEC feature compliance, per Broadcom), provide endpoint candidates that compete with NVIDIA’s supported stack when their driver, transport and collective combination qualifies — the table below maps it. Packet spray re-introduces vendor coupling at the endpoint: spraying packets across all paths only works if the receiving NIC can reassemble out-of-order delivery in hardware, so the switch and NIC must agree — which is why Spectrum-X and UEC are switch+NIC systems, not just switches. The transport semantics that ride on top — lossless vs lossy, in-order vs out-of-order — are the subject of Chapter 8.4.

The back-end SuperNIC field (2026) — physical capacity, Ethernet link structure, transport
NICAggregate BW (physical)Max Ethernet link / breakoutTransport supportAvailability (as of Aug 2026)
NVIDIA ConnectX-7400 Gb/s400GbERoCEv2, IB NDR (shipping)Shipping (Hopper/Blackwell fleets)
NVIDIA ConnectX-8800 Gb/s2×400GbE (single 800GbE link unsupported)RoCEv2, IB XDR 800G, Spectrum-X adaptive RDMA (shipping); MRC (vendor-stated native)Shipping (B300 generation)
NVIDIA ConnectX-9800 Gb/s (C9180; 400G variant exists) — 2 per Rubin GPU = 1.6 Tb/s800GbE (first in family)RoCEv2, IB, Spectrum-X, MRC (vendor-stated)Platform shipment forecast: fall 2026 in NVIDIA’s 2026-05-31 announcement
AMD Pensando Pollara 400400 Gb/s400GbE (1×400 / 2×200 / 4×100)RoCEv2 (shipping); UEC-ready (claim, announced Oct 2024)GA April 2025
AMD Pensando Vulcano800 Gb/s800GbEUEC-ready + MRC (vendor-stated)Helios scale-out NIC; initial Helios shipments FQ3 2026
Broadcom Thor 2400 Gb/s400GbERoCEv2 + hardware congestion control (shipping)Shipping
Broadcom Thor Ultra800 Gb/s800GbEFull UEC feature compliance (claim); MRC 2/4/8-plane, up to 128 paths (documented)Sampling since Oct 2025; no public volume-ship announcement
Physical aggregate bandwidth and maximum Ethernet link speed are different specs — the gap decides whether a NIC attaches as one fat port or several plane-facing ports. Transport entries are labeled: (shipping) = in production fabrics; (documented) = in the vendor's published specs; (claim) = vendor declaration; compare the exact UEC self-attestation edition and interoperation evidence; (announced) = pre-availability. IB = InfiniBand.

DPUs and IPUs: the offload tax and what it buys

A DPU (data processing unit; Intel's term is IPU, infrastructure processing unit) is a NIC with a programmable CPU complex, memory, and accelerators added, sitting in the data path between the host and the wire. The canonical examples are NVIDIA BlueField, AMD Pensando, Intel IPU, and the cloud-captive designs (AWS Nitro, Google's IPU work). The premise is infrastructure offload: move the storage initiator, the encryption, the virtual-network overlay, and the security policy plane off the host x86 and onto the DPU, freeing host cores for the paying workload and creating an isolation boundary the tenant cannot see past.

The 2026 flagship sets the scale of the bet. BlueField-4 pairs a 64-core Arm Neoverse V2 complex (64 billion transistors, ~6× the compute of BlueField-3) with the ConnectX-9 NIC at 800 Gb/s, 128 GB of LPDDR5, a PCIe Gen6 host interface, and an on-board SSD, shipping in 2026 both as a card and integrated into the Vera Rubin NVL72 rack (NVIDIA / HPCwire, 2025). That is a server-class computer on a NIC, which is why the DPU is a decision rather than a default. You are adding a second CPU, a second operating system, and a second control plane to every server. It pays back only if you actually run infrastructure functions on it.

Three offload domains justify a DPU, and each is treated in depth elsewhere in the guide:

  • Storage. The DPU acts as an NVMe-oF initiator and runs the GPUDirect Storage data path, presenting remote flash through the supported storage path into GPU memory (with remote latency and failure behavior still in the budget) and bypassing host CPU data copies on a supported direct path while retaining the required setup and control — and, in the BlueField-4 generation, terminating an Ethernet-attached KV-cache/context-memory tier for inference. The storage data path is built out in Chapter 9.3.
  • Security and isolation. The DPU enforces the tenant VPC overlay, line-rate encryption over RDMA, and microsegmentation policy in hardware the tenant root cannot reach — the hard isolation boundary for multi-tenant GPU clouds. This is the substance of Chapter 11.6 and Chapter 11.7.
  • Virtualization & the overlay. The DPU runs the VXLAN/VPC encapsulation and the software-defined network, so the host hypervisor (or bare-metal stack) is relieved of network virtualization — the model top-tier neoclouds standardize on for bare-metal-with-VPC.
Do you actually need a DPU? — the per-server decision
Deployment contextNIC / SuperNIC aloneAdd a DPU/IPUWhat tips the decision
Single-tenant training clusterUsually sufficient — RDMA transport is in the NICOptional — only if storage/security offload is wantedNo tenant boundary to enforce; host cores often not the bottleneck
Multi-tenant GPU cloud / neocloudDepends on device assignment, IOMMU, RDMA keys and control ownershipSelect when policy must execute outside tenant control and this DPU implements itProve DMA, memory authority, confidentiality and performance isolation separately
Inference fleet with disaggregated KV/storageHost CPU runs the storage initiator — steals coresStrong fit — NVMe-oF + context-memory tier offloadGPUDirect Storage path and KV-cache tier free host cores for serving
Cost-sensitive batch / internal clusterPreferred — fewer control planes to operateHard to justify — added capex + a second OS to patchNo isolation or storage-offload requirement to amortize the DPU
The DPU is a per-server line item across thousands of servers. The right answer depends on what infrastructure functions you have to offload, not on bandwidth alone.

Worked decision: does the DPU recover enough host capacity?

$1,500modeled
Illustrative installed DPU price
Assumed per server; excludes extra lifecycle support.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching input, not vendor quote, tariff or market estimate. Used in the Chapter 8.3 DPU decision.

$250modeled
Illustrative value per recovered core
Assumed over the whole service period.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching input, not vendor quote, tariff or market estimate. Used in the Chapter 8.3 DPU decision.

$0.10/kWhmodeled
Illustrative energy unit price
Assumed flat rate for this sensitivity only.
Sep 2026Guide derivation input — explicitly assumed for the illustrative worked case; not a sourced market price. Operand and procurement replacement requirement are declared in the case.register ↗
Scope & caveats

Guide-selected teaching input, not vendor quote, tariff or market estimate. Used in the Chapter 8.3 DPU decision.

Recovered host budget is 16−6=10 cores; six remaining infrastructure cores pass the 12-core service limit. Multiply ten by the stated value per core, subtract installed price, then subtract 0.060 kW×3×8,760 hours×the stated electricity price. The remaining allowance for added lifecycle/support cost is about $840 per server over three years, rounded after arithmetic. Choose the DPU only if the recovered-core test passes and incremental support stays below that allowance; otherwise retain the host/NIC path or change the offload.

The flip is concrete: support below the unrounded allowance preserves this economic case, while a larger support bill reverses it. A functional boundary that must survive tenant compromise can still require independent enforcement, but that is a separate requirement rather than a fabricated financial saving. Acquire the actual core measurement and support price before release. The GPUDirect Storage design guide supplies the data-path mechanism; Chapter 1.8 owns economic method and Chapter 11.6 owns isolation acceptance.

102.4 Tbps
Tomahawk 6 chip capacity; family production volume announced 2026-03-12
Scope & caveats

102.4 Tb/s chip family; 512×200G or 1,024×100G electrical options. System ports and optical variants require separate support and delivery evidence.

51.2 Tbps
Broadcom Jericho4 deep-buffer router; 3nm; HBM packet memory (~160× on-chip); RoCE over >100 km via 3.2 Tbps HyperPorts
224G-class electrical
Electrical SerDes class; obtain exact interface signaling and FEC
Scope & caveats

Electrical capability class; not an exact selected PMD coded rate, symbol rate, useful payload or optical-fiber count.

up to 1.6 Tb/s
per-GPU scale-out bandwidth on the Rubin platform — two 800 Gb/s ConnectX-9 SuperNICs per GPU, not one 1.6T NIC; double the ConnectX-8 generation; BlueField-4 (integrates a CX-9) at 800 Gb/s
Scope & caveats

1.6 Tb/s is the Rubin per-GPU platform bandwidth (2x800G ConnectX-9), not a single-NIC port: the ConnectX-9 hardware manual lists 800 Gb/s SKUs (C9180 single-OSFP, 800GbE-capable). Rubin partner systems 2H 2026.

64-core
BlueField-4 Arm Neoverse V2 complex (64B transistors, ~6× BF-3 compute); 128 GB LPDDR5; PCIe Gen6
~95%
NVIDIA-reported data throughput on Spectrum-X Ethernet at xAI's 100,000-GPU Colossus, zero flow-collision loss
Scope & caveats

NVIDIA-reported Colossus deployment result under that workload and configuration, not a universal Ethernet, Spectrum-X, or scheduled/VOQ-fabric figure; Colossus is an adaptive-routing/telemetry RoCE fabric, not the scheduled-fabric category. Meta reports tuning RoCE and InfiniBand GenAI clusters to equivalent performance — no common-workload test crowns either transport.

1 per GPU
back-end SuperNIC provisioning (8×400G/800G per 8-GPU server = 3.2–6.4 Tb/s/node) + a separate front-end/storage NIC
2x400GbE
ConnectX-8: the single 800 Gb/s physical port runs Ethernet as 2×400GbE (single 800GbE link unsupported; XDR IB carries 800G on the same port) — the dual-plane substrate
Scope & caveats

Per the official supported-speeds lists (Ethernet max 400GbE per logical link; XDR InfiniBand 800G) and NVIDIA's port-link-type documentation stating a single 800GbE link is unsupported; the B300 reference architecture deploys it as 2x400GbE into two planes.

Oct 14, 2025
Broadcom Thor Ultra: industry-first 800G UEC-compliant AI-Ethernet NIC (packet-level multipathing, OOO placement, selective retransmit) — sampling at announcement
Scope & caveats

'Sampling with select customers' at announcement; no public volume-shipment announcement found as of 2026-08. UEC compliance is Broadcom's product claim; no public certification program exists yet.

Populated-switch power and cooling check

For the same hypothetical leaf, total switch heat is 800 W base electronics +200 W fans/control +64×16 W modules=about 2.0 kW. Four switches contribute about 8.1 kW; the assumed 2.1 kW switch and 9.0 kW rack allowances both pass. These are switch-end modules only: the far-end modules belong to their host or neighboring-switch location. Cold spares add no operating heat, and a supplier total that already includes modules must not receive them twice. Obtain the fully populated input-power curve, cooling-fluid or airflow envelope, fan failure mode and inlet limit; send the location-specific heat to Chapter 5.1 and the equipment qualification to Chapter 5.7. A bare ASIC power number cannot close the rack cooling contract.

Deep dive: collective offload and the supported fallback

A collective offload is both a capability and a resource limit. List supported operations, data types and message sizes; the library path that selects it; concurrent group limits; and the fallback under resource exhaustion or switch failure. An offload premium earns its cost only if the application’s exposed phase improves under the same workload and recovery test. The reduction arithmetic has one home in Chapter 8.6; this silicon contract records the exact supported combination, including the feature lost after a NIC or switch substitution.

The vertical-integration question, restated as silicon

The three ASIC families resolve into one strategic choice: buy the network as a co-designed system from one vendor, or assemble it from merchant silicon. NVIDIA's pitch is that the switch (Quantum/Spectrum), the NIC (ConnectX), and the DPU (BlueField) are designed together, so features like SHARP, packet spray with hardware reassembly, and line-rate encryption work end-to-end out of the box — and NVIDIA’s October 2024 Colossus account reports roughly 95% data throughput for that Spectrum-X deployment, while the purchased combination still needs its own workload and recovery test. The catch is that the system's value depends on owning both ends; the moment you mix in a third-party NIC, the co-designed features degrade or disappear.

The merchant counter-case is that Broadcom (Tomahawk for radix, Jericho for depth) plus a SuperNIC of your choosing plus an open NOS (SONiC, FBOSS) gives you a multi-vendor supply chain, no single-vendor margin capture, and the freedom to mix optics and cables — at the cost of doing the integration yourself and accepting that the most aggressive co-designed features (SHARP-equivalent in-network reduction) are not yet at parity. Choose the merchant Ethernet/RoCE path when a second qualified switch/NIC/NOS combination buys a real price or supply option; reserve the integrated premium for jobs whose supported reduction, routing or recovery features save enough completion time and operating cost to pay for it. The business-model and margin framing of merchant-vs-captive silicon is developed in Chapter 7.1.

Deep dive: why the DPU's second control plane is the part people forget to budget

The DPU sales pitch is all about what it offloads. The part that gets under-budgeted is what it adds: a complete second computer in every server, with its own operating system, its own firmware, its own security-patch cadence, and its own failure modes. When you deploy BlueField at fleet scale, you have doubled the number of OS images you patch, the number of agents you monitor, and the number of things that can break in the data path between the host and the wire — a DPU that hangs can take the server’s network with it when that network’s data path depends on the DPU.

This is why the DPU decision is not "is it powerful" (it obviously is — 64 Arm cores, 800 Gb/s) but "do I have enough infrastructure work to amortize a second control plane across thousands of servers." In a multi-tenant neocloud, the isolation boundary must keep policy outside tenant control. A DPU can enforce that boundary and offload storage, but earns its operational tax only when the implemented functions and measured resource savings justify it; prove device assignment, DMA authority and recovery rather than inferring isolation from a rating. In a single-tenant training cluster with no isolation requirement and host cores to spare, a plain SuperNIC is often the better engineering choice precisely because it is one control plane, not two. The lesson generalizes: every offload frees a host resource and adds an operational surface, and the DPU is only a win when the freed resource is worth more than the added surface. The isolation case that most often tips it is in Chapter 11.6; the storage case in Chapter 9.3.

Anti-patterns

The recurring silicon mis-selections all come from optimizing one number in isolation instead of reading the gating spec and the downstream cost:

  • Buying the aggregate Tbps and ignoring the SerDes generation. A switch bought one SerDes generation behind spends twice the electrical lanes on each port at your speed, so the same aggregate bandwidth arrives as more cages, more front-panel area and a narrower set of qualified port modes and breakouts. Copper runs the other way: a lower per-lane rate reaches further on copper, which is exactly why Broadcom offers Tomahawk 6 as 512×200G and as 1,024×100G. Compare aggregate capacity, lane count, supported port modes, cage count and qualified reach as five separate axes — read the lane, not the headline.
  • Deep buffers everywhere. Putting deep-buffer routing silicon on short-reach leaf switches where a shallow-buffer high-radix ASIC belongs — paying latency, power, and $/port for burst absorption the tier never needs. Match usable buffer to the admitted burst and feedback delay.
  • A DPU with nothing to offload. Deploying BlueField-class silicon across a single-tenant cluster with no isolation requirement and no storage-offload plan — a per-server capex line and a second control plane bought for a feature set you never enable.
  • Mixing a third-party NIC into a co-designed fabric. Buying a Spectrum-X or SHARP-capable switch for its integrated features, then pairing it with a generic NIC that cannot do hardware reassembly or in-network reduction — paying the captive premium and getting the merchant feature set.
This chapter supplies the silicon that the rest of Part 8 wires together. The copper-reach and scale-up domain those ASICs sit inside is Chapter 8.2; the transport semantics (RoCE vs IB, lossless vs lossy, in-order vs out-of-order) that the NICs implement are Chapter 8.4; the radix-and-oversubscription math that turns these chips into a topology is Chapter 8.5; the congestion control, packet spray, and SHARP collective offload that the silicon enables are Chapter 8.6. The DPU's offloaded functions are developed where they live: storage and GPUDirect in Chapter 9.3, multi-tenant isolation in Chapter 11.6, and microsegmentation/zero-trust in Chapter 11.7. The merchant-vs-captive business model behind the family fork is Chapter 7.1; the telemetry that operates all of it is Chapter 10.6 and Chapter 14.2.

Choose the ASIC, NIC and NOS as the combination that closes every required line, including the populated-switch thermal envelope and the feature preserved after substitution. A missing buffer or operation line is grounds to change the configuration or hold the order. Accepting a bigger chip headline instead buys a ceiling that the real ports, software or cooling cannot reach.

Cite this chapter
Fehn, J. (2026). Network Silicon: Switch ASICs, NICs & DPUs (Chapter 8.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-3-network-silicon-switch-asics-nics-and-dpus (accessed 2026-09-29).
@misc{aidc-8-3,
  author       = {Fehn, Jacob},
  title        = {Network Silicon: Switch ASICs, NICs & DPUs (Chapter 8.3)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-3-network-silicon-switch-asics-nics-and-dpus},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit