Chapter 8.3
In this chapter · 9 sections
Network Silicon: Switch ASICs, NICs & DPUs
The switch ASIC, NIC and DPU set the ceiling on every fabric above them, so their SerDes generation, buffer architecture and offload engines must close with the topology before the purchase is frozen.
What you'll decide here
- Which switch ASIC family you standardize on — merchant high-radix Broadcom Tomahawk, merchant deep-buffer Broadcom Jericho, Cisco Silicon One, or NVIDIA Quantum InfiniBand / Spectrum Ethernet — because the supported ASIC + NIC + NOS combination sets radix, usable buffers, queues, telemetry, host attachment, populated-switch heat and vendor lock-in for the cluster.
- Whether your back-end NIC is a plain RoCE/IB NIC or a SuperNIC with full transport offload — and whether the host even needs a DPU, or whether a NIC suffices, because the DPU is a per-server tax you justify with storage, security, and multi-tenant isolation, not with raw bandwidth.
- Shallow shared-buffer vs deep-buffer VOQ at each tier: shallow when usable queues absorb the admitted burst until PFC/ECN or admission acts, deep when the larger bounded burst justifies its power and queueing cost. Crossing buildings or oversubscribing changes the feedback budget, but neither buffer architecture rescues sustained overload.
- Which functions you push off the host CPU and onto the DPU (storage initiator, encryption, the VPC overlay, the security policy plane) versus leaving on x86 — every offloaded function frees host cores for the workload but adds a second control plane to operate and patch.
- Whether the chosen SerDes and port modes carry you through a GPU refresh, because lane rate helps set port count, reach and the copper-vs-optics break. 224G-class SerDes and 200G/lane Ethernet service describe the same interface generation at different boundaries; a higher-rate roadmap becomes useful only with supported MAC modes, FEC and a qualified channel.
The fabric chapters that bracket this one — scale-up in Chapter 8.2, scale-out protocols in Chapter 8.4, topology in Chapter 8.5 — all assume a set of silicon building blocks and reason about how to wire them together. This chapter is about the blocks themselves: the three classes of programmable silicon the entire AI network is assembled from. The switch ASIC moves packets between ports. The NIC (and its beefed-up cousin, the SuperNIC) connects a server's accelerators to the wire and runs the RDMA transport. The DPU/IPU is a NIC with a CPU complex bolted on, sitting in the data path to offload storage, security, and virtualization from the host. Select the silicon well and the fabric design that follows is a series of well-posed wiring problems; under-spec the SerDes, mismatch the usable buffer allocation to the admitted burst, or buy a DPU you have no offload for, and you have baked a ceiling into the cluster that no topology cleverness can lift.
These decisions are unusually consequential because they come early and stay sticky. The switch ASIC family fixes your radix and your lock-in posture before a single cable is run. The SerDes generation on that ASIC fixes, for the next several years, how many ports you get, how far copper reaches, and when optics become mandatory. And the DPU decision — whether to deploy one at all, and what to run on it — is a per-server line item across thousands of servers that pays back only if you actually move work onto it.
The gating spec: SerDes generation, not aggregate Tbps
Marketing leads with “102.4 Tbps switch” or “800G NIC,” but the gating specification is the complete port contract: MAC service rate, electrical lanes, coded signaling, FEC, cage arrangement and supported breakout. The lane still governs how many interfaces fit and how far the electrical route reaches before optics add power and cost. A vendor’s 224 Gb/s-class SerDes label and 200 Gb/s-per-lane Ethernet service are not consecutive generations. Multiplying the electrical ceiling by lane count does not produce useful Ethernet throughput. Cisco’s G300 data sheet dated 2026-02-10 lists those electrical and MAC specifications separately. IEEE P802.3dj remains a draft project; specify the interface document and FEC mode rather than declaring a product port to be a ratified application.
Why does the lane matter alongside the aggregate? Broadcom’s Tomahawk 6 launch illustrates the slicing: its 2025 announced 102.4 Tb/s MAC capacity supports 64 × 1.6T, 128 × 800G or 256 × 400G logical-port modes, while a purchased box implements only some of those arrangements. An optical module can also gearbox between different electrical and optical lane counts. Count ASIC-to-module lanes, optical transmit lanes and active fibers separately; none is the number of faceplate cages. Read supported modes from the switch and module release together, then confirm the host/NIC path reaches them. The copper cliff remains a route constraint, not a universal meter count: lower-rate electrical lanes can improve a qualified copper route while consuming more lanes per port, so “newer SerDes” is not a complete reach or power decision. The physical-layer consequences belong in Chapter 8.9; the transport operations implemented at the endpoints belong in Chapter 8.4.
Switch ASIC families: the supported system behind the silicon fork
Switch silicon differs on buffer organization, forwarding features and support boundary as well as aggregate rate. Broadcom Tomahawk and Jericho, Cisco Silicon One G300, and NVIDIA Spectrum/Quantum all belong in the same comparison. Buffering philosophy decides where bursts wait, while the business model decides who controls support and replacements. Separate the ASIC supplier from the system supplier and the NOS: a chip’s SDK feature does not prove that a purchased software release exposes it. Freeze one supported combination of SKU, firmware, driver, NOS, transceiver and collective library, then ask which row prevents that combination from carrying the topology. A missing queue-allocation or telemetry line is a contract gap even when every headline bandwidth number passes.
Broadcom Tomahawk is the merchant high-radix line: it favors on-chip shared buffering and endpoint-controlled traffic over storing prolonged overload in external memory. Broadcom’s 2026-03-12 release says the 102.4 Tb/s Tomahawk 6 family is shipping in production volume. That record follows the initial launch; it does not prove volume availability for every package or optical variant. PFC/ECN and endpoint admission still have to limit bursts before the shared queues fill. Demand the actual port map and the selected NOS’s queue, load-balancing and counter support before turning chip radix into leaf capacity.
Broadcom Jericho is the merchant deep-buffer, VOQ line: routing-class silicon that spends external memory on bursts too large for an on-chip queue budget. Broadcom’s August 2025 Jericho4 announcement advertises a 51.2 Tb/s family aggregate and HBM-backed packet buffering; those resources buy time for a bounded arrival/service mismatch, not more sustained capacity at the destination link. Compare packet-routing mode with scheduled-fabric mode, including fabric elements, admission protocol, power and backlog tail. Longer routes enlarge the feedback budget, but even a deep buffer eventually fills under sustained overload. Chapter 8.8 admits cross-site traffic against the protected path.
NVIDIA Quantum and Spectrum provide InfiniBand and Ethernet system paths. Quantum-X800’s 800G/port InfiniBand with SHARPv4 in-network reduction — 14.4 TFLOPS of in-network compute, nine times the prior generation (NVIDIA, March 2024; mechanics in Chapter 8.6) — and Spectrum-X paired with ConnectX/BlueField endpoints are concrete integrated candidates in the dated NVIDIA records. Their collective offload and endpoint/switch features have to survive the exact software combination being purchased. Cisco G300 supplies another current Ethernet candidate on the same axes. A controlled workload comparison can justify an integrated premium; NVIDIA’s October 2024 Colossus report of roughly 95% data throughput is one deployment result, not a transport winner. The economic choice is whether one support boundary saves enough integration and recovery effort to outweigh the loss of interchangeable components. Chapter 7.1 develops the merchant-versus-captive business model.
| Contract line | Tomahawk 6 | Cisco G300 | Spectrum-6 class | Jericho class | Quantum-X800 |
|---|---|---|---|---|---|
| MAC capacity / modes | 102.4 Tb/s family; system must expose selected modes | 102.4 Tb/s; 64×1.6T in data sheet | 102.4 Tb/s ASIC in SN6000 manual | Named SKU and forwarding mode required | Named InfiniBand switch port configuration required |
| Electrical lanes | 512×200G or 1,024×100G options | 512×224G-class SerDes; separate from MAC rate | Use selected SN6000 electrical port mode | Exact SKU lane map required | Exact switch/NIC link mode required |
| Cages / optics | System and package-specific; qualify each populated mode | System cage map required; ASIC radix is not cage count | SN6600-LD manual: 64 OSFP cages; mode-specific ports | Line card and fabric-module map required | Physical cages and logical ports counted separately |
| Buffer / pool | Shared on-chip; obtain usable allocation | Fully shared on-die; allocation depends on policy | Manual: 160 MB per ASIC shared; verify usable allocation | VOQ/external memory by SKU and system mode | Credit-managed resources; obtain per-class limits |
| Queues / classes | NOS class mapping, limits and pool policy required | Multiple output queues; selected NOS limits required | NOS class mapping and pool policy required | VOQ destination/class map and credits required | Virtual lanes, arbitration and management policy required |
| Path selection / reorder | Document enabled routing; endpoint reorder gates spray | ECMP/weighted, flow/flowlet, spray; endpoint contract | Integrated routing depends on NIC/firmware combination | Scheduled or packet mode; state which | Subnet-manager/routing configuration required |
| Telemetry | Counters, microburst capture and timestamp origin required | ECN, PFC watchdog, flow/queue and event telemetry listed | Expose port/queue/error and selected transport counters | VOQ depth, grants, fabric health and endpoint progress | Port/error, route and offload/fallback counters |
| NOS / support | System vendor and exact NOS release | SDK/SAI and SONiC reference listed; system release required | Named NVIDIA/system software support matrix | Interface/fabric elements and software as one contract | Switch, NIC, manager and collective release as one contract |
| Evidence date / status | Broadcom volume announcement 2026-03-12 | Cisco data sheet 2026-02-10; no job uplift inferred | SN6000 manual, preliminary May 2026; supply by SKU | Obtain dated selected-system evidence | Obtain dated selected-system evidence |
Scope & caveats
Chip specifications, not guaranteed system cages, usable buffers or application acceleration. Electrical SerDes class differs from MAC service rate.
Scope & caveats
SN6000 manual revision 1.0, May 2026 (publication day unspecified; date normalized to month start). SN6600-LD has 64 OSFP cages with mode-specific logical ports; no shipment-date inference.
Shallow-shared vs deep-buffer VOQ: the buffering tradeoff
Tomahawk versus Jericho is the cleanest hardware fork in networking, and it has no free lunch. The question it turns on: where do you store a packet that arrives faster than its egress port can drain? Two silicon philosophies answer it differently.
Shallow shared-buffer ASICs (Tomahawk-class) keep a small, fast pool of on-chip SRAM shared across all ports. The bet is that with short reach, a non-blocking topology, and good endpoint congestion control (PFC/ECN/DCQCN, adaptive routing, packet spray — the machinery of Chapter 8.6), bursts are absorbed at the source and the switch never needs to hold much. The payoff is lowest latency, highest radix, and lowest power-per-bit. The risk: when a real incast burst exceeds the shallow buffer, you must either drop (lossy) or assert backpressure (PFC), and PFC at scale brings head-of-line blocking and deadlock risk. Shallow buffering only works if the congestion-control loop is fast enough to keep the buffer from filling.
Deep-buffer VOQ ASICs (Jericho-class) attach large off-chip memory — HBM on Jericho4 — and organize it as virtual output queues: a separate logical queue per egress destination, so a congested port cannot head-of-line-block traffic bound elsewhere. The payoff is the ability to absorb enormous bursts and to run lossless over long reach — a larger burst envelope for RoCE over long DCI routes when admission and feedback keep arriving bytes within it; depth alone cannot prevent drops. The cost is real: added latency (a packet may sit in deep buffer), higher power and die area (HBM is not free), and higher $/port. You do not want deep buffers on a short-reach leaf where they add latency you never needed; you do want them at the fabric edge that crosses buildings.
Close each line before choosing the leaf. The MAC total is 64×800 Gb/s=51.2 Tb/s per direction; downstream and upstream each carry 32×800=25.6 Tb/s. The cage map consumes 64×4=256 electrical lanes, while the optical map has 64×8=512 transmit lanes and 64×16=1,024 active fibers. Those counts are distinct. A 102.4 Tb/s ASIC can have enough aggregate capacity without the selected box exposing these cages or modes. A dual-400GbE-only adapter fails the single-800GbE attachment line even if its aggregate is 800 Gb/s.
The required usable queue allocation is 64×0.60=38.4 MB, inside the assumed 40 MB pool, but candidate A’s per-port cap of 0.512 MB fails. Reject that configuration before comparing throughput. The flip is a documented and tested per-port allocation of at least 0.60 MB while all 64 required queues are active and the other pool reservations still fit. At that threshold the arithmetic closes; purchase remains HOLD until the cage map, host payload, telemetry and concurrent-buffer test produce evidence. The G300 data sheet and SN6000 manual illustrate why advertised chip features and system allocations are separate. Chapter 13.7 accepts the installed combination.
NICs and SuperNICs: the RoCE/IB offload path
The NIC is where the network meets the accelerator, and in AI fabrics it does far more than push frames. The defining feature is RDMA — remote direct memory access — which lets a GPU on one node read or write registered GPU memory on another without host data copies when registration, PCIe/IOMMU topology and software support the path (with the CPUs retaining the required setup and control), the foundation that makes collective communication tolerable at scale. Two transports carry RDMA: native InfiniBand (NVIDIA ConnectX in IB mode) and RoCEv2 (RDMA over Converged Ethernet), which runs the same verbs over a routable Ethernet/UDP underlay. The NIC implements the transport in hardware; the quality of that implementation — how it handles congestion, retransmission, and packet reordering — is part of why one fabric advances the job while another wastes bisection capacity on retries or idle barrier wait; measure the selected job’s service rate and completion tail.
One supported back-end reference pattern is one NIC attachment per GPU: an 8-GPU server carries 8×400G or 8×800G back-end ports (3.2–6.4 Tb/s/node), plus a separate, smaller NIC for the front-end/storage/management plane. The term SuperNIC denotes the AI-optimized variant: full transport offload, hardware support for the adaptive routing / packet-spray and out-of-order reassembly that Ultra Ethernet and Spectrum-X require, and line-rate congestion handling. ConnectX-8 carries the current 800 Gb/s generation — with a port-structure subtlety that shapes fabric design: its single 800 Gb/s physical port runs Ethernet as 2×400GbE (a single 800GbE link is unsupported; the same port does carry 800 Gb/s XDR InfiniBand), which is why the B300-generation reference design wires each GPU into two independent 400GbE planes (NVIDIA, 2025–26; → the planes pattern in Chapter 8.5). The Rubin platform doubles per-GPU scale-out bandwidth to 1.6 Tb/s, delivered as two 800 Gb/s ConnectX-9 SuperNICs per GPU — eight per four-GPU compute tray — the 800 Gb/s C9180-class parts being the family's first with single-port 800GbE (a 400 Gb/s variant also exists; NVIDIA, 2026). The pair teaches the general rule: read a NIC's supported link speeds alongside its aggregate, because the two differ across the whole 2026 field. And the field is no longer single-vendor: AMD's Pensando Pollara 400 (announced October 2024 as the industry's first UEC-ready AI NIC, generally available April 2025) and 800G Vulcano, and Broadcom's Thor 2 and 800G Thor Ultra (full UEC feature compliance, per Broadcom), provide endpoint candidates that compete with NVIDIA’s supported stack when their driver, transport and collective combination qualifies — the table below maps it. Packet spray re-introduces vendor coupling at the endpoint: spraying packets across all paths only works if the receiving NIC can reassemble out-of-order delivery in hardware, so the switch and NIC must agree — which is why Spectrum-X and UEC are switch+NIC systems, not just switches. The transport semantics that ride on top — lossless vs lossy, in-order vs out-of-order — are the subject of Chapter 8.4.
| NIC | Aggregate BW (physical) | Max Ethernet link / breakout | Transport support | Availability (as of Aug 2026) |
|---|---|---|---|---|
| NVIDIA ConnectX-7 | 400 Gb/s | 400GbE | RoCEv2, IB NDR (shipping) | Shipping (Hopper/Blackwell fleets) |
| NVIDIA ConnectX-8 | 800 Gb/s | 2×400GbE (single 800GbE link unsupported) | RoCEv2, IB XDR 800G, Spectrum-X adaptive RDMA (shipping); MRC (vendor-stated native) | Shipping (B300 generation) |
| NVIDIA ConnectX-9 | 800 Gb/s (C9180; 400G variant exists) — 2 per Rubin GPU = 1.6 Tb/s | 800GbE (first in family) | RoCEv2, IB, Spectrum-X, MRC (vendor-stated) | Platform shipment forecast: fall 2026 in NVIDIA’s 2026-05-31 announcement |
| AMD Pensando Pollara 400 | 400 Gb/s | 400GbE (1×400 / 2×200 / 4×100) | RoCEv2 (shipping); UEC-ready (claim, announced Oct 2024) | GA April 2025 |
| AMD Pensando Vulcano | 800 Gb/s | 800GbE | UEC-ready + MRC (vendor-stated) | Helios scale-out NIC; initial Helios shipments FQ3 2026 |
| Broadcom Thor 2 | 400 Gb/s | 400GbE | RoCEv2 + hardware congestion control (shipping) | Shipping |
| Broadcom Thor Ultra | 800 Gb/s | 800GbE | Full UEC feature compliance (claim); MRC 2/4/8-plane, up to 128 paths (documented) | Sampling since Oct 2025; no public volume-ship announcement |
DPUs and IPUs: the offload tax and what it buys
A DPU (data processing unit; Intel's term is IPU, infrastructure processing unit) is a NIC with a programmable CPU complex, memory, and accelerators added, sitting in the data path between the host and the wire. The canonical examples are NVIDIA BlueField, AMD Pensando, Intel IPU, and the cloud-captive designs (AWS Nitro, Google's IPU work). The premise is infrastructure offload: move the storage initiator, the encryption, the virtual-network overlay, and the security policy plane off the host x86 and onto the DPU, freeing host cores for the paying workload and creating an isolation boundary the tenant cannot see past.
The 2026 flagship sets the scale of the bet. BlueField-4 pairs a 64-core Arm Neoverse V2 complex (64 billion transistors, ~6× the compute of BlueField-3) with the ConnectX-9 NIC at 800 Gb/s, 128 GB of LPDDR5, a PCIe Gen6 host interface, and an on-board SSD, shipping in 2026 both as a card and integrated into the Vera Rubin NVL72 rack (NVIDIA / HPCwire, 2025). That is a server-class computer on a NIC, which is why the DPU is a decision rather than a default. You are adding a second CPU, a second operating system, and a second control plane to every server. It pays back only if you actually run infrastructure functions on it.
Three offload domains justify a DPU, and each is treated in depth elsewhere in the guide:
- Storage. The DPU acts as an NVMe-oF initiator and runs the GPUDirect Storage data path, presenting remote flash through the supported storage path into GPU memory (with remote latency and failure behavior still in the budget) and bypassing host CPU data copies on a supported direct path while retaining the required setup and control — and, in the BlueField-4 generation, terminating an Ethernet-attached KV-cache/context-memory tier for inference. The storage data path is built out in Chapter 9.3.
- Security and isolation. The DPU enforces the tenant VPC overlay, line-rate encryption over RDMA, and microsegmentation policy in hardware the tenant root cannot reach — the hard isolation boundary for multi-tenant GPU clouds. This is the substance of Chapter 11.6 and Chapter 11.7.
- Virtualization & the overlay. The DPU runs the VXLAN/VPC encapsulation and the software-defined network, so the host hypervisor (or bare-metal stack) is relieved of network virtualization — the model top-tier neoclouds standardize on for bare-metal-with-VPC.
| Deployment context | NIC / SuperNIC alone | Add a DPU/IPU | What tips the decision |
|---|---|---|---|
| Single-tenant training cluster | Usually sufficient — RDMA transport is in the NIC | Optional — only if storage/security offload is wanted | No tenant boundary to enforce; host cores often not the bottleneck |
| Multi-tenant GPU cloud / neocloud | Depends on device assignment, IOMMU, RDMA keys and control ownership | Select when policy must execute outside tenant control and this DPU implements it | Prove DMA, memory authority, confidentiality and performance isolation separately |
| Inference fleet with disaggregated KV/storage | Host CPU runs the storage initiator — steals cores | Strong fit — NVMe-oF + context-memory tier offload | GPUDirect Storage path and KV-cache tier free host cores for serving |
| Cost-sensitive batch / internal cluster | Preferred — fewer control planes to operate | Hard to justify — added capex + a second OS to patch | No isolation or storage-offload requirement to amortize the DPU |
Worked decision: does the DPU recover enough host capacity?
Scope & caveats
Guide-selected teaching input, not vendor quote, tariff or market estimate. Used in the Chapter 8.3 DPU decision.
Scope & caveats
Guide-selected teaching input, not vendor quote, tariff or market estimate. Used in the Chapter 8.3 DPU decision.
Scope & caveats
Guide-selected teaching input, not vendor quote, tariff or market estimate. Used in the Chapter 8.3 DPU decision.
Recovered host budget is 16−6=10 cores; six remaining infrastructure cores pass the 12-core service limit. Multiply ten by the stated value per core, subtract installed price, then subtract 0.060 kW×3×8,760 hours×the stated electricity price. The remaining allowance for added lifecycle/support cost is about $840 per server over three years, rounded after arithmetic. Choose the DPU only if the recovered-core test passes and incremental support stays below that allowance; otherwise retain the host/NIC path or change the offload.
The flip is concrete: support below the unrounded allowance preserves this economic case, while a larger support bill reverses it. A functional boundary that must survive tenant compromise can still require independent enforcement, but that is a separate requirement rather than a fabricated financial saving. Acquire the actual core measurement and support price before release. The GPUDirect Storage design guide supplies the data-path mechanism; Chapter 1.8 owns economic method and Chapter 11.6 owns isolation acceptance.
Scope & caveats
102.4 Tb/s chip family; 512×200G or 1,024×100G electrical options. System ports and optical variants require separate support and delivery evidence.
Scope & caveats
Electrical capability class; not an exact selected PMD coded rate, symbol rate, useful payload or optical-fiber count.
Scope & caveats
1.6 Tb/s is the Rubin per-GPU platform bandwidth (2x800G ConnectX-9), not a single-NIC port: the ConnectX-9 hardware manual lists 800 Gb/s SKUs (C9180 single-OSFP, 800GbE-capable). Rubin partner systems 2H 2026.
Scope & caveats
NVIDIA-reported Colossus deployment result under that workload and configuration, not a universal Ethernet, Spectrum-X, or scheduled/VOQ-fabric figure; Colossus is an adaptive-routing/telemetry RoCE fabric, not the scheduled-fabric category. Meta reports tuning RoCE and InfiniBand GenAI clusters to equivalent performance — no common-workload test crowns either transport.
Scope & caveats
Per the official supported-speeds lists (Ethernet max 400GbE per logical link; XDR InfiniBand 800G) and NVIDIA's port-link-type documentation stating a single 800GbE link is unsupported; the B300 reference architecture deploys it as 2x400GbE into two planes.
Scope & caveats
'Sampling with select customers' at announcement; no public volume-shipment announcement found as of 2026-08. UEC compliance is Broadcom's product claim; no public certification program exists yet.
Populated-switch power and cooling check
For the same hypothetical leaf, total switch heat is 800 W base electronics +200 W fans/control +64×16 W modules=about 2.0 kW. Four switches contribute about 8.1 kW; the assumed 2.1 kW switch and 9.0 kW rack allowances both pass. These are switch-end modules only: the far-end modules belong to their host or neighboring-switch location. Cold spares add no operating heat, and a supplier total that already includes modules must not receive them twice. Obtain the fully populated input-power curve, cooling-fluid or airflow envelope, fan failure mode and inlet limit; send the location-specific heat to Chapter 5.1 and the equipment qualification to Chapter 5.7. A bare ASIC power number cannot close the rack cooling contract.
Deep dive: collective offload and the supported fallback
A collective offload is both a capability and a resource limit. List supported operations, data types and message sizes; the library path that selects it; concurrent group limits; and the fallback under resource exhaustion or switch failure. An offload premium earns its cost only if the application’s exposed phase improves under the same workload and recovery test. The reduction arithmetic has one home in Chapter 8.6; this silicon contract records the exact supported combination, including the feature lost after a NIC or switch substitution.
The vertical-integration question, restated as silicon
The three ASIC families resolve into one strategic choice: buy the network as a co-designed system from one vendor, or assemble it from merchant silicon. NVIDIA's pitch is that the switch (Quantum/Spectrum), the NIC (ConnectX), and the DPU (BlueField) are designed together, so features like SHARP, packet spray with hardware reassembly, and line-rate encryption work end-to-end out of the box — and NVIDIA’s October 2024 Colossus account reports roughly 95% data throughput for that Spectrum-X deployment, while the purchased combination still needs its own workload and recovery test. The catch is that the system's value depends on owning both ends; the moment you mix in a third-party NIC, the co-designed features degrade or disappear.
The merchant counter-case is that Broadcom (Tomahawk for radix, Jericho for depth) plus a SuperNIC of your choosing plus an open NOS (SONiC, FBOSS) gives you a multi-vendor supply chain, no single-vendor margin capture, and the freedom to mix optics and cables — at the cost of doing the integration yourself and accepting that the most aggressive co-designed features (SHARP-equivalent in-network reduction) are not yet at parity. Choose the merchant Ethernet/RoCE path when a second qualified switch/NIC/NOS combination buys a real price or supply option; reserve the integrated premium for jobs whose supported reduction, routing or recovery features save enough completion time and operating cost to pay for it. The business-model and margin framing of merchant-vs-captive silicon is developed in Chapter 7.1.
Deep dive: why the DPU's second control plane is the part people forget to budget
The DPU sales pitch is all about what it offloads. The part that gets under-budgeted is what it adds: a complete second computer in every server, with its own operating system, its own firmware, its own security-patch cadence, and its own failure modes. When you deploy BlueField at fleet scale, you have doubled the number of OS images you patch, the number of agents you monitor, and the number of things that can break in the data path between the host and the wire — a DPU that hangs can take the server’s network with it when that network’s data path depends on the DPU.
This is why the DPU decision is not "is it powerful" (it obviously is — 64 Arm cores, 800 Gb/s) but "do I have enough infrastructure work to amortize a second control plane across thousands of servers." In a multi-tenant neocloud, the isolation boundary must keep policy outside tenant control. A DPU can enforce that boundary and offload storage, but earns its operational tax only when the implemented functions and measured resource savings justify it; prove device assignment, DMA authority and recovery rather than inferring isolation from a rating. In a single-tenant training cluster with no isolation requirement and host cores to spare, a plain SuperNIC is often the better engineering choice precisely because it is one control plane, not two. The lesson generalizes: every offload frees a host resource and adds an operational surface, and the DPU is only a win when the freed resource is worth more than the added surface. The isolation case that most often tips it is in Chapter 11.6; the storage case in Chapter 9.3.
Anti-patterns
The recurring silicon mis-selections all come from optimizing one number in isolation instead of reading the gating spec and the downstream cost:
- Buying the aggregate Tbps and ignoring the SerDes generation. A switch bought one SerDes generation behind spends twice the electrical lanes on each port at your speed, so the same aggregate bandwidth arrives as more cages, more front-panel area and a narrower set of qualified port modes and breakouts. Copper runs the other way: a lower per-lane rate reaches further on copper, which is exactly why Broadcom offers Tomahawk 6 as 512×200G and as 1,024×100G. Compare aggregate capacity, lane count, supported port modes, cage count and qualified reach as five separate axes — read the lane, not the headline.
- Deep buffers everywhere. Putting deep-buffer routing silicon on short-reach leaf switches where a shallow-buffer high-radix ASIC belongs — paying latency, power, and $/port for burst absorption the tier never needs. Match usable buffer to the admitted burst and feedback delay.
- A DPU with nothing to offload. Deploying BlueField-class silicon across a single-tenant cluster with no isolation requirement and no storage-offload plan — a per-server capex line and a second control plane bought for a feature set you never enable.
- Mixing a third-party NIC into a co-designed fabric. Buying a Spectrum-X or SHARP-capable switch for its integrated features, then pairing it with a generic NIC that cannot do hardware reassembly or in-network reduction — paying the captive premium and getting the merchant feature set.
Choose the ASIC, NIC and NOS as the combination that closes every required line, including the populated-switch thermal envelope and the feature preserved after substitution. A missing buffer or operation line is grounds to change the configuration or hold the order. Accepting a bigger chip headline instead buys a ceiling that the real ports, software or cooling cannot reach.
Cite this chapter
Fehn, J. (2026). Network Silicon: Switch ASICs, NICs & DPUs (Chapter 8.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-3-network-silicon-switch-asics-nics-and-dpus (accessed 2026-09-29).
@misc{aidc-8-3,
author = {Fehn, Jacob},
title = {Network Silicon: Switch ASICs, NICs & DPUs (Chapter 8.3)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-3-network-silicon-switch-asics-nics-and-dpus},
note = {Accessed 2026-09-29}
}