The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 7.2

In this chapter · 6 sections
Term help

NVIDIA Accelerators: Hopper → Blackwell → Vera Rubin → Rubin Ultra → Feynman

NVIDIA publishes an annual platform roadmap; the purchase can include the chip, rack or pod, and each boundary changes the power, cooling, fabric and refresh work you inherit.

POWER-BOUNDDENSITY-RAMPGOODPUT

What you'll decide here

  1. Which named Hopper, Blackwell or Rubin system and operating mode to procure, and whether its complete power, cooling, physical and software profile fits the site.
  2. Whether your unit of purchase is the GPU, the HGX board, or the rack-scale NVL system (NVL72/144/576) — because the scale-up domain you buy is the scale-up domain you are stuck with until refresh.
  3. Whether to ride the annual cadence at every step (Hopper → Blackwell → Blackwell Ultra → Rubin → Rubin Ultra → Feynman) or skip generations — and how the choice performs across the contested 2–3-year obsolescence bear case and published 4–6-year useful-life estimates.
  4. For inference at long context, whether to adopt disaggregated serving (Groq LPX decode racks + Rubin GPUs for prefill/attention) or stay monolithic — a fork that changes your BOM, your fabric, and your cost per token.
  5. Whether to reserve busbar, floor-loading and water capacity for NVIDIA’s announced Kyber / 800 VDC / roughly 600 kW planning case (H2 2027 roadmap, checked September 2026), or accept the later enabling-work cost when that rack will not fit the hall.
HPE GB200, Lenovo GB300, Pegatron Vera Rubin and the announced Kyber option retain their own power and availability boundaries. A different rack can require floor, power and cooling work that a rack swap cannot provide.

NVIDIA sells a cadence: an annual rhythm of accelerator generations, each one re-drawing the rack, the fabric, and the power chain underneath it. The question it forces is not which GPU is fastest but which rung of the ramp you are committing your building to — and what the next rung costs if you guessed wrong. Since 2022 the unit of purchase has migrated upward: from the H100 as a board, to the HPE GB200 NVL72 profile as a 132 kW rack you buy whole, to the Vera Rubin NVL72 and the Rubin Ultra NVL576 / Kyber NVL144 generation as multi-rack pods plumbed for 800 VDC. Each migration can commit the facility to a new interface; a roadmap announcement is neither a delivery promise nor an accepted rack. You can defer the silicon; you cannot defer the floor loading, the water, and the interconnection slot the silicon implies.

The roadmap here reads as a sequence of decisions and their downstream costs. The per-GPU specs run across Hopper → Blackwell → Blackwell Ultra → Vera Rubin → Rubin Ultra → Feynman; the NVL system — not the GPU — became the unit of procurement once the size of the scale-up domain you buy (8 → 72 → 144 → 576 GPUs) turned into a strategic commitment rather than a datasheet line; the disaggregated-inference fork that Rubin CPX opened is now answered by Groq's LPU/LPX decode racks; and the annual roadmap cadence compresses competitors' design windows and can pressure your economic refresh schedule without changing book life. Per-GPU NVLink bandwidth appears here as a datasheet attribute; the NVLink/NVSwitch fabric that aggregates it has its canonical home in Chapter 8.2.

You are buying a power envelope, not a FLOPS number

The instinct is to compare generations on peak FLOPS. That is the marketing-number trap (Chapter 7.1): the headline figures are sparse FP4 with all the asterisks stripped, and they tell you almost nothing about what you must build. The number that actually cascades through your facility is rack power. An H100 air-cooled rack lands near 40 kW; the HPE GB200 NVL72 profile is rated 132 kW and mandates direct-to-chip liquid; a GB300 NVL72 carries an up-to-142 kW facility design basis; the Vera Rubin VR200 NVL72 carries a 330 kW cabinet facility design basis; and Rubin Ultra's Kyber rack (NVL144) targets ~600 kW on an 800 VDC bus. Those are different OEM operating profiles, facility allowances and roadmaps, not a 15× measured power trajectory.

A hall scoped for the previous generation needs an interface-by-interface retrofit check. You do not get to "upgrade" from a 40 kW air hall to a 132 kW liquid hall by swapping racks — the floor loading must be checked against the selected OEM’s populated operating mass and point loads, the airflow, electrical headroom and facility water must each satisfy the selected profile. The density wall (Chapter 5.1) and the DLC default (Chapter 5.4) are downstream of which rung of this ramp you bought into. The power curve governs; the compute curve follows.

Hopper → Blackwell → Vera Rubin → Rubin Ultra → Feynman: the per-GPU arc

Hopper (H100, 2022 / H200, 2024) is the generation most of the installed base still runs. H100 ships 80 GB HBM3 at ~3.35 TB/s, ~700 W TDP, FP8 Transformer Engine, NVLink 4 at 900 GB/s per GPU. H200 is the same compute die with 141 GB HBM3E at ~4.8 TB/s — a memory-bandwidth refresh that disproportionately helps inference decode. Hopper is air-coolable, which is exactly why it became the default and why the jump to Blackwell broke so many facility assumptions.

Blackwell (B200 / GB200, 2024–2025) is a dual-die GPU — two reticle-limited compute dies on one package behaving as a single CUDA device over a 10 TB/s die-to-die link — with 180 GB of usable HBM3E (HGX B200; NVIDIA's current GB200 specification gives 372 GB across the superchip's two GPUs — 186 GB usable per Blackwell GPU and ~13.4 TB across an NVL72, against the 192 GB raw figure earlier material quoted), a second-generation Transformer Engine adding native FP4, and NVLink 5 at 1.8 TB/s per GPU. The GB200 superchip pairs two Blackwell GPUs with one Grace CPU over NVLink-C2C. Blackwell Ultra (B300 / GB300, 2025) lifts HBM to 288 GB and adds steady-power and transient-mitigation features (capacitor energy storage, ramp smoothing) that exist because a ~135 kW-TDP / 155 kW-peak rack toggling between idle and full all-reduce is a grid problem (Chapter 7.12).

Vera Rubin (H2 2026) is the next platform, not just a chip — "VR200" is supply-chain/roadmap shorthand, not NVIDIA's spelled-out name. The Rubin GPU is again dual-die on a 4-reticle CoWoS-L interposer — ~336 billion transistors, 1.6x Blackwell — with 288 GB HBM4 across 8 stacks at up to ~22 TB/s, sixth-generation Tensor Cores, and NVLink 6 at 3.6 TB/s per GPU. The Vera CPU is NVIDIA's custom Arm successor to Grace. The rack-scale unit is the NVL72 — 72 Rubin GPUs and 36 Vera CPUs (NVIDIA briefly marketed it as "NVL144" by counting the 144 compute dies inside the 72 dual-die packages, then reverted to the package-based NVL72 name at CES/GTC 2026), delivering ~3.3x the GB300 NVL72 on inference, ~3.6 EF FP4 inference / ~1.2 EF FP8 training per rack, with ~260 TB/s of scale-up NVLink bandwidth. NVIDIA says the Vera Rubin platform is in full production and partner availability is planned for the second half of 2026.

Rubin Ultra (H2 2027) is where the unit of purchase jumps again. As shown at GTC 2026 it packs four compute dies per package (~100 PFLOPS FP4, 1 TB HBM4e per package) and deploys two ways: the ~600 kW Kyber rack — NVL144: 144 quad-die packages = 576 GPU compute dies on 800 VDC, ~15 EF FP4 inference / ~5 EF FP8 training, with eight Kyber racks forming NVL1152 — and the NVL576 system, eight MGX NVL72-class racks stitched into one 576-GPU optical NVLink domain. (Supply-chain reports from Apr–Jun 2026 say Rubin Ultra was scaled back to a dual-die "2+2" Kyber board layout over CoWoS-L warpage and yield — unconfirmed by NVIDIA, so the per-package quad-die specs are in flux; the ~600 kW rack-level power and H2 2027 timeline are unaffected.) Feynman (2028) is the next architecture on the roadmap — paired with the new Rosa CPU, advanced 3D die stacking, and NVLink 8, with the ConnectX/Spectrum generations advancing in lockstep; it is widely expected (though not confirmed by NVIDIA) to move to a sub-2 nm TSMC node with backside power delivery. The public roadmap targets an approximately annual platform cadence; dates, product configuration and volume availability remain roadmap commitments until shipped and independently verified.

NVIDIA accelerator generations — the per-GPU and per-rack arc
GenerationGPU memoryMem BWNVLink/GPU (bidirectional aggregate)TDP/GPURack unitRack powerAvailability
Hopper H10080 GB HBM3~3.35 TB/s900 GB/s (NVLink 4)~700 WHGX 8-GPU / DGX~40 kW (air)2022
Hopper H200141 GB HBM3E~4.8 TB/s900 GB/s (NVLink 4)~700 WHGX 8-GPU~40 kW (air)2024
Blackwell GB200186 GB usable HBM3E~8 TB/s1.8 TB/s (NVLink 5)~1.0–1.2 kWNVL72 rack132 kW nominal (HPE; DLC)2024–2025
Blackwell Ultra GB300288 GB HBM3E~8 TB/s1.8 TB/s (NVLink 5)~1.4 kWNVL72 rack135 kW TDP / 155 kW peak (Lenovo; DLC)2025
Vera Rubin NVL72288 GB HBM4~22 TB/s3.6 TB/s (NVLink 6)~1.8 kW (Max Q) / ~2.3 kW (Max P)NVL72 rackPegatron RA4803-72N3: 188 kW Max Q / 228 kW Max P2026 — partner availability planned H2 2026; first measured silicon Jul 2026 (CoreWeave)
Rubin Ultra1 TB HBM4e/pkg (announced; spec in flux Aug 2026)(4-die pkg)(NVLink 7)not publishedKyber NVL144 (rack) / NVL576 (8-rack)~600 kW (800 VDC)H2 2027 (announced)
FeynmanHBM4e+ (roadmap; configuration not published)Not published(NVLink 8)Not publishedKyber-classNo contracted profile2028 (roadmap)
Per-GPU figures are NVIDIA datasheet / roadmap; 2026+ rows are announced/early-shipping — the first Vera Rubin NVL72 rack was validated Jun 2026 (CoreWeave/Dell) and the first measured silicon result followed 2026-07: ~10x tokens/s per MW vs GB200 NVL72 on DeepSeek R1 at matched interactivity (CoreWeave, all NVIDIA software optimizations on) — NVIDIA says full production; partner availability planned H2 2026. Rubin Ultra and Feynman remain roadmap, and Rubin Ultra's memory spec is in flux (TrendForce 2026-08-04: 8-Hi HBM4e / 12-Hi HBM4 / 8-Hi HBM4 under evaluation vs the 12-Hi HBM4e baseline; NVIDIA unconfirmed). Pegatron's RA4803-72N3 datasheet specifies Vera Rubin rack power at 188 kW (Max Q) / 228 kW (Max P); NVIDIA DSX sets a 330 kW cabinet TDP facility design basis. FLOPS are peak FP4 dense; see Chapter 7.1 for the dense-vs-sparse, peak-vs-sustained discount. Rack-power and HBM figures cross-checked against the guide's Numbers Register (Appendix D). For Vera Rubin, ~1.8 kW at Max Q and ~2.3 kW at Max P are operating profiles for the same package; Rubin Ultra has no published per-package TDP.
186 GBderived
GB200 usable HBM per GPU
Scope & caveats

372/2=186 GB; raw 192 GB is a different capacity boundary.

The rack-power column governs the table, not the FLOPS column. The compute numbers grow impressively, but they are the easy part — their delivery and packaging schedule is itself a material risk. What strands capital is the rightmost columns: the rack unit changes shape (board → 72-GPU rack → 144 → 576), the power per rack escalates an order of magnitude, and the cooling and voltage architecture flip underneath. Moving between generations requires a gap analysis across rack geometry, mass/reactions, power topology and transients, coolant/TCS/FWS conditions, residual air, fabric, controls, safety and service routes. Reserved capacity can reduce work, but neither a one-decision upgrade nor demolition is universal. The generation you choose is the substrate you commit to.

The Rubin line crossed from roadmap to silicon in mid-2026: CoreWeave brought up and validated the industry-first Vera Rubin NVL72 rack on June 1, 2026 (Livingston, NJ; Dell PowerEdge XE9812-integrated) — 72 Rubin GPUs plus 36 Vera CPUs, ~3.6 EFLOPS FP4, 260 TB/s NVLink 6 all-to-all — and its power basis is now published: Pegatron's RA4803-72N3 specifies 188 kW (Max Q) / 228 kW (Max P), while NVIDIA DSX sets a 330 kW cabinet TDP facility design basis. Size irreversible infrastructure to the facility design basis; run energy and TCO models on the chosen operating profile.

Why the NVL system became the unit of purchase

Through Hopper, the unit was the 8-GPU HGX board and the scale-up domain was 8 GPUs wide. Blackwell broke that model: the GB200 NVL72 fuses 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain — 18 compute trays and 9 NVSwitch trays connected by a copper NVLink spine carrying ~130 TB/s of aggregate rack bandwidth across more than 5,000 in-rack copper cables — so that every GPU has load/store addressability over NVLink to ~13.4 TB of distributed HBM across the domain; remote access does not make every GPU's memory local, uniform-latency or universally cache-coherent (~1.44 EF FP4 sparse). The integrated rack is the procurement and acceptance boundary, including when an OEM assembles it from qualified parts. The reason this matters strategically is that the scale-up domain size is now a purchasing decision that sets your parallelism ceilings until your next refresh.

The consequence runs in both directions. A wide domain (72 → 144 → 576 GPUs) lets you fit tensor-parallel and pipeline-parallel groups, and especially wide expert-parallel MoE inference, entirely inside the NVLink fabric — where per-direction bandwidth is roughly 9x the scale-out NIC on a current 800G port and 18x on a 400G one — instead of spilling collectives onto the slower back-end network. Wide-EP MoE serving (e.g., EP32 vs EP8) is the canonical workload that the big domain unlocks (Chapter 8.2). But a wide domain you do not use is stranded capital: a latency-bound 8B-parameter inference service pinned to a 72-GPU NVLink rack is paying for a fabric it never lights up. Buy the smallest supported domain that the working set and measured collectives require — the NVL system is a commitment, not a default.

Deep dive: NVLink per-GPU bandwidth as a datasheet attribute (and where the fabric lives)

Every generation advertises a per-GPU NVLink number — 900 GB/s on Hopper (NVLink 4), 1.8 TB/s on Blackwell (NVLink 5), 3.6 TB/s on Rubin (NVLink 6) — and it is tempting to treat it like memory bandwidth, a property of the chip. It is not. The per-GPU figure is the bidirectional aggregate injection bandwidth into a switched fabric — transmit and receive summed, so halve it before comparing against any one-way number — and what you actually get depends on the NVSwitch generation, the domain size, and the topology that aggregates it. On an NVL72 the 1.8 TB/s per GPU aggregates to ~130 TB/s of rack scale-up bandwidth; on Rubin NVL72 the 3.6 TB/s per GPU aggregates to ~260 TB/s per rack; the 40-rack Vera Rubin POD (1,152 Rubin GPUs — 40 racks across five rack types, only ~16 of them compute racks) reaches ~10 PB/s of aggregate scale-up bandwidth (NVIDIA's POD-level figure, counting all rack types). The datasheet attribute is comparable across vendors (it is roughly an order of magnitude above the scale-out NIC), but the design decisions it drives — switch-tray count, copper-vs-optical reach, NVLink-SHARP in-network reduction, domain partitioning — are fabric decisions.

So we record the per-GPU number here, in the accelerator chapter, because it is a property you compare when choosing silicon. We engineer the fabric that consumes it — NVSwitch topology, NVLink-SHARP collective offload, the copper-reach wall that is pushing Rubin Ultra toward optical scale-up, and how the scale-up domain is partitioned and scheduled — in Chapter 8.2. Treat the two as a split: the chip chapter owns the attribute; the network chapter owns the system.

The disaggregated-inference fork: Rubin CPX, and the Groq LPU/LPX decode path

Rubin introduces a second, quieter fork that reshapes the inference BOM. Long-context inference has two phases with opposite hardware profiles: the context (prefill) phase reads and encodes the entire input — compute-bound, hungry for FLOPS, light on memory bandwidth — while the generation (decode) phase emits tokens one at a time, memory-bandwidth-bound and latency-sensitive, leaning on HBM and the KV cache. A monolithic GPU sized for decode (expensive HBM) is overpaying to do prefill; a GPU sized for prefill is starved on memory for decode. NVIDIA's original answer was Rubin CPX, announced 2025-09-09 — a context-phase accelerator with ~30 PFLOPS NVFP4, 3x attention acceleration over GB300, and 128 GB of GDDR7 rather than HBM, roughly 5x more cost-effective per byte than HBM for this compute-bound role.

CPX's roadmap status is now uncertain. Credible secondary reporting — The Next Platform quoting NVIDIA's Ian Buck, corroborated by Tom's Hardware and The Elec — says CPX was reportedly removed from the roadmap at GTC 2026 (March), before production silicon and about six months after its September 2025 announcement. NVIDIA issued no official cancellation and has not reaffirmed CPX on a current roadmap page, so treat it as announced, roadmap status uncertain / reportedly pulled — not as shipping, and not as formally cancelled. In the reported reshuffle the context/prefill role does not move to a dedicated chip: Rubin GPUs handle prefill and attention themselves, and the latency-critical decode phase is reassigned to Groq silicon under NVIDIA's ~$20B Groq deal.

The two Groq products are easy to conflate, so name them precisely. The chip is the Groq 3 LPU — an SRAM-based inference processor with ~500 MB of on-chip SRAM and ~150 TB/s of SRAM bandwidth, no HBM, that cannot compute prefill attention. The rack is the LPX: 256 LPUs with ~128 GB of aggregate on-chip SRAM, the deployable unit that runs the token-by-token decode loop (the FFN/MoE layers) while the KV cache and attention stay on the Rubin HBM GPUs. (One secondary outlet speculates a CPX-style context accelerator could resurface with Feynman around 2028 — speculation, not an NVIDIA statement.)

A commercial-structure footnote, not a silicon change (Aug 2026): the post-acquisition Groq entity announced it is an NVIDIA Cloud Partner (Groq, 2026-08-12) — buying B300/GB300/Rubin GPUs alongside LPUs and renting the capacity out, with SemiAnalysis noting many of the clusters Groq rents are straight Blackwell with no LPUs at all. Groq-as-neocloud now coexists with the LPU decode path described above; nothing in NVIDIA's CPX/LPU integration story has officially changed.

The decision: for million-token-context workloads, do you adopt disaggregated serving — a pool of Rubin (HBM) GPUs doing prefill and attention feeding a pool of SRAM-based Groq LPX decode racks doing low-latency generation (FFN/MoE), coupled over the fabric by streaming intermediate activations between the attention and FFN stages — or stay monolithic? The KV cache stays on the Rubin GPUs, so this is not the KV-handoff transport of a classic prefill–decode split. Disaggregation wins on cost-per-token and interactive latency at long context because you run the token-by-token decode on Groq's cheap, deterministic SRAM silicon instead of monopolizing HBM GPUs for generation; it costs you a more complex serving stack (separate pools, KV-cache transport, careful ratio tuning) and a fabric that must move KV cache between phases efficiently (Chapter 10.11). For short-context, latency-flat workloads the disaggregation overhead does not pay. But the fork is wider than long context: NVIDIA describes prefill–decode, attention–FFN and external-drafter arrangements, and each imposes a different traffic volume and latency requirement — name the arrangement before you derive a fabric requirement from it.

132 kW nominal TDP: 115 kW liquid + 17 kW air
NVIDIA GB200 NVL72 by HPE: 132 kW nominal rack TDP, with 115 kW liquid and 17 kW air heat-removal duties
Scope & caveats

Exact HPE product profile. Keep this 132 kW / 115 kW liquid / 17 kW air record separate from the OCP MGX Rev. 1.1 reference profile of 120 kW / approximately 102 kW liquid / 18 kW air.

~600 kWforecast
Rubin Ultra Kyber rack (NVL144) on 800 VDC — GTC 2026 also defined NVL576 as an eight-rack MGX system; supply-chain reports (Apr–Jun 2026) say the package was cut to dual-die 2+2 over CoWoS-L warpage, unconfirmed by NVIDIA
Scope & caveats

NVIDIA's published figure (GTC 2025) is 600 kW per Rubin Ultra Kyber rack and GTC 2026 did not revise it. SemiAnalysis (2026-05-26) reports Kyber Ultra 'approaching 660 kW' — a single-source analyst estimate for a 2027 part, recorded here rather than adopted, since the vendor primary figure still stands.

3.6 TB/s
NVLink 6 per-GPU bandwidth (Rubin); 900 GB/s Hopper, 1.8 TB/s Blackwell — ~260 TB/s per NVL72 rack
~288 GB
HBM4 per Rubin GPU at ~22 TB/s; trajectory H100 80 GB → H200 141 GB → HGX B200 180 GB → B300 288 GB, flat into Rubin, then Rubin Ultra 1 TB
Scope & caveats

2.75 TB/s = Rubin config (22 TB/s ÷ 8× 36 GB stacks). JEDEC HBM4 base ~2.0 TB/s/stack; shipping 2026 HBM4 parts run ~2.56 (SK hynix) → ~2.8 (Micron) → up to ~3.3 (Samsung) TB/s/stack.

~336 B
transistors per Rubin GPU (dual-die, 4-reticle CoWoS-L) — 1.6x Blackwell's ~208 B
GTC 2026 (reported)
Groq 3 LPU chip (~500 MB SRAM, ~150 TB/s) → LPX rack (256 LPUs, ~128 GB aggregate SRAM) fills the disaggregated decode/FFN slot (no HBM); Rubin GPUs retain prefill + attention. Rubin CPX reportedly pulled at GTC 2026 (no official NVIDIA cancellation)
Scope & caveats

Rubin CPX reportedly pulled (GTC 2026, no official NVIDIA cancellation); decode reassigned to Groq LPU chip / LPX rack under the ~$20B deal

Status per credible secondary reporting; no official NVIDIA cancellation or roadmap reaffirmation as of 2026-07-21.

annual
architecture cadence — Blackwell 2024, Blackwell Ultra 2025, Rubin 2026, Rubin Ultra 2027, Feynman 2028
2–3 yr (bear case) vs 4–6 yr (GS)forecast
accelerated GPU economic life — contested: bear case 2–3 yr on obsolescence; Goldman Sachs estimates 4–6 yr useful life vs 5–6 yr book
Scope & caveats

Goldman Sachs characterizes four to six years as the estimated useful life of AI accelerators; it does not substantiate a general two-to-three-year economic life or a universal 20–40% three-year residual value.

188 kW (Max Q) / 228 kW (Max P)
Vera Rubin NVL72 rack power — 188 kW Max Q / 228 kW Max P (Pegatron); 330 kW cabinet TDP as the facility design basis (NVIDIA DSX)
Scope & caveats

Operating profiles from the OEM rack datasheet, which also ships 4 × 110 kW power shelves. Distinct from the facility design basis: NVIDIA's DSX Facilities Infrastructure Design Guide v2.0 (2026-08-19) provisions a 330 kW cabinet TDP for Vera Rubin NVL72. Size substrate, busway, pipe and CDU from the design basis; run energy and TCO models on Max Q / Max P.

TrendForce (2026-06-25) independently put VR200 rack draw at ~225 kW, inside the OEM Max Q–Max P band. NVIDIA publishes no rack-power figure on the product page.

1st rack validated 2026-06-01
CoreWeave industry-first Vera Rubin NVL72 bring-up (Jun 1 2026, Livingston NJ; Dell-integrated) — 72 Rubin GPUs + 36 Vera CPUs, ~3.6 EFLOPS FP4, 260 TB/s NVLink 6; a single validated rack, not fleet GA
Scope & caveats

Single validated rack; fleet general availability not yet. Dell factory-integration ('under 6.5 hours' delivery-to-production) per Dell; ~3.6 EFLOPS FP4 rack figure per DataCentre Magazine.

The annual cadence as a strategic weapon

The cadence is a competitive instrument as much as a delivery schedule, and it cuts two ways. Against competitors, a yearly architecture compresses the window any challenger has to close a gap: if AMD or a custom-ASIC roadmap matches Blackwell as a qualified Rubin system arrives, the comparison resets (Chapter 7.3, Chapter 7.5). The software moat (CUDA, the NCCL/Dynamo/TensorRT stack) compounds this — a one-year hardware tick gives the ecosystem a fresh target every twelve months. Against the buyer, the same cadence is a depreciation accelerant: a frontier accelerator's economic life is tested at 2–3 years in the guide's contested bear case, against published 4–6-year useful-life estimates and 5–6-year book policies, because next year's part does the same work at materially lower cost-per-token. The cadence that protects NVIDIA's lead also shortens your amortization runway.

That sets up the buyer's real decision — ride every generation, or skip? Riding each step can improve qualified output-per-dollar but means continuous capital outlay and the operational churn of new power, cooling, and fabric envelopes every year. Skipping a generation (e.g., Hopper → Rubin, bypassing Blackwell) reduces churn and lets one substrate investment serve longer, at the cost of running a generation behind on token economics during the gap. The deciding variables are your residual-value assumption (do used GPUs hold enough value to backstop the refresh — see Chapter 1.8) and whether your facility substrate can even accept the generation you would skip to. A 40 kW air-cooled position cannot accept a roadmap 600 kW Kyber profile without facility work. An announcement offers a refresh option: exercise it only when Chapter 1.8 covers hardware, migration, facility work, downtime and the value of keeping the fleet.

Ride-every-generation vs skip-a-generation — the refresh fork
StrategyToken-economics positionCapital cadenceSubstrate/ops churnBest fit
Ride every generationAlways at the frontier of cost-per-tokenContinuous, annual outlayHigh — new power/cooling/fabric envelope each yearFrontier labs; neoclouds competing on price/token
Skip one generationOne step behind during the gapLumpy, every ~2 yearsModerate — one substrate serves two cyclesEnterprises with stable workloads; substrate-constrained sites
Hold (run to economic end)Falls behind; relies on residual demandMinimal until forced refreshLowest — until a hard substrate/density wallBatch/offline inference; depreciation-sensitive operators
Heuristic, not a rule; the right answer is set by residual-value assumptions, substrate readiness, and token-economics sensitivity. Chapter 7.11 qualifies eligible equipment; Chapter 1.8 owns cost denominators, earning-life sensitivity and NPV.
Deep dive: the NVL72 vs NVL144 naming and the dual-die accounting trap

The naming will trip up anyone reading the roadmap as a procurement spec. GB200 NVL72 means 72 Blackwell GPUs in the NVLink domain — and each of those GPUs is itself a dual-die package, so the rack contains 144 compute dies behaving as 72 CUDA devices. Vera Rubin NVL72 keeps the same package-based accounting — 72 Rubin GPUs (dual-die packages), the same 72-package footprint as GB200 NVL72, not double. (NVIDIA briefly marketed this rack as "NVL144" by counting the 144 compute dies inside those 72 dual-die packages, then reverted to the package-based NVL72 name at CES/GTC 2026, so older slides showing "NVL144" refer to the same rack.) Rubin Ultra's Kyber NVL144 then means 144 quad-die packages = 576 compute dies per rack — a genuine 2x in package count over NVL72 and a 4x in die count — while NVL576 now names the eight-rack MGX system (576 packages in one optical domain), not a single rack.

Why this matters in practice: if you size power, cooling, and fabric by reading "144" as "twice the GPUs of 72," you will mis-budget. The honest comparison is die-to-die and rack-to-rack power: NVL72 at ~132 kW, Rubin NVL72 at 188 kW Max Q / 228 kW Max P (Pegatron; same 72-package footprint, with ~1.8 kW / ~2.3 kW per package at those operating profiles), with a 330 kW cabinet facility design basis (NVIDIA DSX), the Kyber NVL144 rack at ~600 kW (double the packages, an 800 VDC rack). Always reduce the marketing nomenclature to (packages per rack) × (per-package power) before you put a number in a design-basis document — multiply by dies per package only when the last factor is power per die, or you double-count the very thing the naming trap is about. On the figures above, 72 dual-die packages at ~1.8 kW each is ~130 kW of accelerator power, and accelerator power is not rack input power until you add CPUs, NICs, switch trays and conversion losses for one named OEM configuration. The cross-vendor version of this discipline — dense vs sparse, peak vs sustained, the marketing-number trap — is the subject of Chapter 7.1.

What the roadmap commits, and what it leaves reversible

A defensible accelerator strategy uses the same discipline that governs facility scoping (Chapter 1.1): sort the decisions by the cost of changing your mind. The roadmap makes some things reversible and some irreversible, and they are not the ones people assume.

  • Reversible (defer, re-decide at refresh): the specific accelerator generation within a power/cooling envelope — a hall plumbed for ~140 kW liquid still needs the exact GB200 or GB300 electrical, liquid, room-air and service profile checked; the HGX-vs-NVL choice for an inference fleet; the disaggregation decision for inference serving; the ride-vs-skip refresh cadence.
  • Irreversible (commit at scoping): whether the hall is plumbed for liquid at all; the floor-loading basis for 3,000-lb-plus wet racks; the electrical capacity and voltage path (415/480 VAC vs an 800 VDC future); reserved physical headroom — busbar runs, pipe racks, switchgear room — for the next density step; and Rubin readiness, because VR200 NVL72 at a 330 kW cabinet facility design basis is a substrate step rather than a rack swap. Price the option premium now against later retrofit cost; a reservation is not automatically cheap.

The strategic move is the same as everywhere in this guide: convert irreversible decisions into reversible ones while the option is cheap. Reserve the busbar capacity and water for a density step-up you have not committed to; buy into the NVL domain your workload uses today but provision the substrate for the domain you might buy in two years. Design the rungs you can reach into the building before you need them; Chapter 7.13 records manufacturer, revision, mode, interfaces and enabling works, and Chapter 1.8 prices the refresh.

The accelerator landscape and the datasheet-reading discipline that frames this chapter are in Chapter 7.1; the open challenger and hyperscaler XPUs that the annual cadence is aimed at are in Chapter 7.3 and Chapter 7.4; the HBM and packaging constraints that gate every generation are in Chapter 7.6 and Chapter 7.7; on-package power delivery and transient mitigation in Chapter 7.12; the rack-as-integration-unit in Chapter 7.13; and the qualified configuration from Chapter 7.11 enters the lifecycle cost model in Chapter 1.8 for TCO and refresh economics. The NVLink/NVSwitch fabric that consumes the per-GPU bandwidth recorded here is engineered in Chapter 8.2; scale-out topology and oversubscription in Chapter 8.5; disaggregated inference serving in Chapter 10.11. The density wall and DLC default the rack-power ramp forces are in Chapter 5.1 and Chapter 5.4; the 800 VDC transition in Chapter 4.7; and the reversible-vs-irreversible scoping discipline in Chapter 1.1.
Cite this chapter
Fehn, J. (2026). NVIDIA Accelerators: Hopper → Blackwell → Vera Rubin → Rubin Ultra → Feynman (Chapter 7.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-2-nvidia-accelerators-hopper-blackwell-vera-rubin-rubin-ultra-feynman (accessed 2026-09-29).
@misc{aidc-7-2,
  author       = {Fehn, Jacob},
  title        = {NVIDIA Accelerators: Hopper → Blackwell → Vera Rubin → Rubin Ultra → Feynman (Chapter 7.2)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-2-nvidia-accelerators-hopper-blackwell-vera-rubin-rubin-ultra-feynman},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit