The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Appendix B

In this chapter · 4 sections
Term help

Reference Designs & Worked Examples

Choose the accelerator and counted reference population, then reconcile rack power, supporting IT, flow, fiber, ports and quoted dollars at their own boundaries; this appendix integrates two reference builds.

What you'll decide here

  1. Start from the per-archetype design-basis sheet, then freeze the named product: density, cooling, fabric blocking, redundancy and GPU:CPU/storage/network ratios require the workload and service tests in Chapters 1.7, 7.13 and 8.5.
  2. Treat the scalable unit (SU) as the atomic rack purchase and deployment block: multiply its rack population, then recount shared CDUs, fabric tiers, storage and supporting IT at the new campus or cluster size.
  3. Use the 50 MW campus and 100k-GPU BOM as order-of-magnitude calibrators: counts follow the stated ratios; the August 2025 server estimate, 2026 assumed allowances and unquoted networking have distinct purchase boundaries. Obtain project quotes before releasing capital.
  4. When a vendor proposal disagrees on racks, CDUs, switches or optics, reconcile the count before signing: oversubscription, redundancy, generation and purchase scope explain differences; no universal percentage tolerance makes a missing core acceptable.
  5. Re-derive, do not interpolate, when you change generation (GB200 → GB300 → Vera Rubin → Kyber): density, flow, and busbar current step discontinuously, so the multipliers in §2 are generation-stamped on purpose.

This appendix is the reusable arithmetic layer behind Part 1's archetype framework (Chapter 1.1) and Part 1's requirements matrix (Chapter 1.7). It does four things, in order: (1) a per-archetype design-basis sheet that freezes the inputs every later number inherits; (2) a scalable-unit (SU) budget giving the power, cooling, water, and network draw of one atomic deployment block per accelerator generation; (3) a 50 MW campus sized from the SU up, with the power chain, cooling plant, and water loop derived; and (4) a 100k-GPU cluster reference BOM with counts and costs for GPUs, racks, CDUs, switches, optics and storage: an August 2025 GB200 estimate, 2026 assumptions and unquoted networking.

The method throughout is multiplier-first. A reference design is a chain of ratios: GPUs per rack, racks per SU, kW per rack, L/min per kW, NICs per node, optics per NIC, GB/s per GPU of storage. Once those are pinned, rack aggregates are multiplications you can audit; shared plant and fabric tiers need the explicit ledgers below. Counts in the tables are exact arithmetic from the stated ratios. Dated server estimates and assumed allowances need project quotes before budget release. Density, flow, and current figures are generation-stamped because they step discontinuously across GB200 → GB300 → Vera Rubin → Rubin Ultra Kyber; do not interpolate across a generation boundary.

1. Per-archetype design-basis sheets

The design-basis sheet is the single page that everything downstream inherits — the concrete instantiation of the workload-profile and design-basis artifacts named in Chapter 1.1. Choose the column that matches your dominant archetype, and the rest of the appendix is parameterized for you. The two reference builds in §3–§4 use the frontier-training column unless noted, because it is the most constraining; an inference-shaped build has its own basis: enterprise HGX relaxes density and fabric, while frontier serving runs the same NVL72-class racks and often buys more service redundancy, not less.

Design-basis sheet by workload archetype (2026 reference points)
ParameterFrontier trainingPost-training / RLOnline inference (frontier / enterprise)Batch inferenceEdge inference
Dominant acceleratorGB200/GB300 NVL72Disaggregated: NVL72 trainer + HGX rolloutFrontier: GB300 NVL72; enterprise: HGX B300 / RTX PROHGX B200; prior-gen acceptableL4/L40S, Jetson, single B200
Named NVL72 rack quantitiesHPE GB200: 132 kW nominal; Lenovo GB300: 135 kW nominal / up to 155 kW workload-dependent peakTrainer: HPE GB200 132 kW nominalFrontier: Lenovo GB300 135 kW nominal / up to 155 kW workload-dependent peakSelect batch product separatelySelect edge product separately
Illustrative workload rack densitiesUse the named training product aboveRollout: assumed 40–60 kW HGXEnterprise: assumed 40–60 kW HGXBatch: assumed 30–60 kWNetworked edge DC: assumed 5–50 kW
Cooling modalityDLC mandatory, warm-waterDLC trainer + RDHx/air rolloutFrontier: DLC mandatory; enterprise: air or RDHxAir often sufficientAir / sealed modular
Scale-up domain72 GPUs (→144, →576)72 trainer / 8 rollout72 (frontier NVL72); 8 (enterprise HGX node)81 (single node)
Scale-out fabricMeasured collectives + step-time targetDerive trainer and rollout tiers separatelyMeasured request/KV/EP traffic + tail SLOMeasured throughput traffic + completion targetMeasured local traffic + WAN/SLO boundary
Fabric transportInfiniBand XDR or Spectrum-XIB trainer / RoCE rolloutEthernet/RoCE commonEthernet, cost-optimizedStandard IP
GPU:CPU ratio2:1 (NVL72: 72G:36C)2:1 trainer / 4–8:1 rollout2:1 frontier NVL72; 4:1–8:1 enterprise8:1+1:1 appliance
GPU:storage (BW)Derive loader, checkpoint drain and restore independently (9.4/9.8)Trainer like trainingKV-cache tier and prefill/decode disaggregation; model load tierStreaming object tierLocal NVMe only
Resilience inputs (not topology)Checkpoint/restart loss; maintenance/fault states; recovery SLOTrainer/rollout state models; staleness and recovery limitsServing SLO; replica/zone/region capacity; site-state continuityQueue/backlog and completion limits; recovery contractLatency routing, backhaul dependence, fleet correlation and recovery
Electrical design basisNamed OEM peak + facility basis — never a TDP multiplierPer named trainer / rollout productsFrontier: named OEM facility basis; enterprise: nameplate + measured profileNamed product figuresNamed product figures
Siting driverCheap firm MW + cold climateFollows dominant sub-workloadSub-50 ms to usersCheapest / curtailable MWLatency budget (30/50/100 ms)
Density and fabric figures are GB200/GB300 NVL72-class, 2026-current. GPU:CPU and GPU:storage are design ratios, not hard limits. The online-inference column is two tiers: frontier serving runs the same liquid-cooled NVL72-class racks as training, and air-cooled 40–60 kW HGX is the enterprise pattern, not the frontier one. See keynumbers for sources and vintages.
192 kW printed EDPp; 1.5 × 132 kW = 198 kW
HPE GB200 rack EDPp and multiplier conflict; electrical release awaits OEM clarification
Scope & caveats

HPE GB200 only; printed peak and recommended busway quantity, not a continuous facility duty. Obtain OEM clarification of the inconsistent multiplier and peak duration.

2. The scalable unit (SU): power / cooling / water / network budget

The scalable unit is the atomic deployment and costing block — order it, integrate it at the factory (L11/L12), ship it, energize it, repeat. Sizing the SU once and then multiplying is what makes campus and cluster arithmetic tractable. We anchor the SU to the NVIDIA DGX SuperPOD GB200 reference: 8 × NVL72 racks = 576 GPUs per SU, with the full SuperPOD at 16 SUs (128 racks, 9,216 GPUs). The budget below gives one SU's draw across four generations; later sections count SUs, not racks.

The cross-generation columns exist because the multipliers step. A GB200 SU is ~1.06 MW of IT; the same 8-rack SU at Kyber density (~600 kW/rack) is ~4.8 MW — a 4.5× jump in what eight racks draw; qualify the Kyber rack's dimensions, mass, and service clearances before assuming it lands on the same footprint. This is the density-ramp trap from Chapter 1.1 as a budget line: the floor, water, and busbar you reserve today must survive it.

Scalable-unit budget — 8-rack SU, by accelerator generation (GPU counts are packages; die counts noted)
MetricHPE GB200 NVL72 (2026 receipt)Lenovo NVIDIA GB300 NVL72 (2025–26)Vera Rubin: Pegatron power / NVIDIA facility basisKyber NVL144 (H2 2027)
GPUs per SU576576576 (72/rack; 1,152 dies)1,152 (144/rack; 4,608 dies)
Rack power (TDP / Max Q)132 kW135 kW nominal TDP (Lenovo LP2357, August 30, 2026)188 kW Max Q~600 kW (facility planning point; TDP not published)
IT power per SU (TDP / Max Q)~1.06 MW1.08 MW~1.50 MW (Max Q)~4.8 MW (planning point)
IT power per SU (peak / facility basis)HPE peak/busway: dated record above; multiplier conflict requires OEM clarificationUp to ~1.24 MW (8 × Lenovo’s up-to-155 kW peak); obtain Lenovo project electrical design requirements~1.82 MW Max P / ~2.64 MW facility basis (330 kW/cabinet)not published; size to the ~600 kW/rack planning point
DLC heat to liquid~0.92 MW~0.97 MW (assumed 90%)~1.50 MW (assumed 100% liquid, Max Q)~4.8 MW (assumed 100% liquid)
Residual air heat~0.14 MW (~17 kW/rack)~0.11 MW (assumed 10%)~0 (assumed 100% liquid)~0 (assumed 100% liquid)
Secondary-loop flow (stated water properties, 10.0 K rise)~1,320 L/min~1,400 L/min~2,160 L/min~6,900 L/min
Coolant inlet / ΔT targetFWS target W45; HPE publishes no technology-loop inlet maximum for this rack / 7–10 °C design rise; the 10.0 K used below is its upper endLenovo LP2357: up to 45 °C inlet; 10 K here is an assumed heat-balance riseAssumed ~45 °C FWS target; qualify Pegatron TCSAssumed ~45 °C FWS target; qualify future TCS
Illustrative back-end NIC ports (one 400G attachment/GPU)576 assumed portsAcquire purchased port profileAcquire purchased port profileAcquire purchased port profile
Illustrative leaf host-facing ports576 assumed portsFollow purchased port profileFollow purchased port profileFollow purchased port profile
Illustrative host-channel modules only1,152 ends; inter-switch ends counted belowSelect modules from purchased port profileSelect modules from purchased port profileSelect modules from purchased port profile
Storage service per SUAcquire loader and checkpoint budgets (9.4/9.8)Acquire workload-specific budgetsAcquire workload-specific budgetsAcquire workload-specific budgets
SU = 8 NVL72-class racks = 576 GPUs (DGX SuperPOD GB200 SU definition). Secondary-loop flow uses the stated water properties and 10.0 K rise: 60,000/(998 × 4.182 × 10.0) L/min per kW of liquid heat, not total rack TDP, before display rounding; recalculate for the approved fluid and operating point. Facility-water make-up assumes evaporative rejection, while dry, non-evaporative rejection needs only loop fill and maintenance top-up. The network rows are a generic 400G host-attachment scenario from Chapter 8.5, separate from each OEM power column; inter-switch links are counted in the campus/cluster port map below. Coolant-inlet cells state the facility-water (FWS) class the design targets; the approved technology-loop (TCS) inlet and return come from the named rack and CDU documentation — an ASHRAE W class does not certify a product setpoint.
Scalable-unit inputs for the two reference builds

Take the GB200 column and walk it forward so the multipliers are explicit. GPUs: 8 racks × 72 GPUs = 576. IT power (TDP): 8 × 132 kW = 1.056 MW ≈ 1.06 MW. Electrical chain: use HPE’s facility guidance and Chapter 4.5’s transient budget, not a TDP multiplier. Heat split: at ~115 kW liquid + ~17 kW air per rack, liquid carries 8 × 115 = 920 kW and air carries 8 × 17 ≈ 140 kW (unrounded nominal sum: 920 + 136 = 1,056 kW). Flow illustration: keep the 920 kW liquid heat boundary fixed. With the stated water properties and 10.0 K rise, 920 kW × 60,000 / (998 × 4.182 × 10.0) ≈ 1,320 L/min for the SU (about 165 L/min per rack). Recalculate with the approved fluid, operating temperature and selected ΔT; do not use the 1,056 kW total rack TDP as TCS heat. Network inputs: use Chapter 8.5’s generic one-port-per-GPU 400G ledger for the two maps below. They are illustrative, not HPE/NVIDIA qualification; inter-switch optics do not scale by a fixed per-GPU multiplier. → SU definition in Chapter 1.7; fabric sizing in Chapter 8.5.

3. Worked example: a 50 MW campus sized from the SU up

Now multiply. The brief: a 50 MW-class IT frontier-training reference campus on HPE GB200 NVL72, built as integer SUs, with the power chain, cooling plant, and water loop derived. We size the IT budget on rack TDP and the electrical chain on each platform's published peak and OEM facility design basis (HPE’s printed GB200 peak conflicts with its stated multiplier; obtain the corrected duty and duration schedule), and we reserve floor/water/busbar headroom for a GB300 → Vera Rubin density step (the irreversible substrate from Chapter 1.1).

SU count. Choose 48 SUs as a clean reference arrangement: three 16-SU SuperPOD-scale halls, or six 8-SU halls. That is 48 × 8 = 384 NVL72 racks and 48 × 576 = 27,648 GPUs. Treating 132 kW/rack as an exact nominal cap fixture, total IT is 50.688 MW of GPU-rack IT, which exceeds a hard 50.000 MW cap by 0.688 MW. If 50.000 MW is a hard cap, use at most 47 SUs: 376 racks, 27,072 GPUs, and 49.632 MW of GPU-rack IT. Supporting network, storage and management IT still have to fit: Chapter 4.1 closes the total-load ledger; Chapter 5.1 carries every residual-air and liquid duty into the plant. Neither rack count proves total-IT admission.

50.688 MW reference campus — 48 SUs (384 racks, 27,648 GB200 GPUs)
SubsystemSizing basisQuantity / value
Scalable unitsReference arrangement48 SUs
NVL72 racks48 × 8384 racks
GPUs384 × 7227,648 GPUs
GPU-rack IT power (nominal TDP)384 × 132 kW50.688 MW (48-SU reference)
Facility power (assumed PUE ≈ 1.2)50.688 MW × 1.2~61 MW
Utility interconnect (N, assumed 15% margin)50.688 MW × 1.2 × 1.15 (round only the result)~70 MW POI / 2× 132 kV feeders
Main transformers≥2 × 75 MVA (N+1 at MV)2–3 × 75 MVA
MV distribution33/13.8 kV ring or radialper-hall 13.8 kV → 415 V / 800 VDC
Ride-through stack (protected-IT boundary)50.688 MW continuous nominal IT; any declared rack peak/transient is allocated across rack capacitance/BBU and upstream UPS/BESSDo not treat a sub-second rack peak multiplier as central-UPS continuous MW
Backup generation (generator-terminal critical-load boundary)50.688 MW IT × 1.2 PUE ≈ 61 MW steady, plus explicit derating/operating reserve and auxiliaries outside that boundary~65–75 MW is a screening installed-capacity band only if the load schedule proves it; ride-through bridges rack-side sub-second transients
DLC heat to facility water384 × 115 kW~44.2 MW thermal
CDUs (HPE 1.3 MW; maximum eight racks/CDU; local 3+1 per 24-rack pod)16 pods × (3 duty + 1 valved standby); each duty CDU = 8 × 115 kW = 0.920 MW64 installed (48 duty + 16 standby)
Secondary-loop flow44,160 kW × 60,000 / (998 × 4.182 × 10.0)~63,500 L/min aggregate (~165 L/min/rack)
Heat rejection~61 MW total heattowers/dry-coolers + adiabatic, economized
Water make-up (assumed WUE ~0.5 L/kWh of IT energy)50.688 MW × 1,000 kW/MW × 0.5 L/kWh × 8,760 h/year~200 ML/year (about 600 m³/day average; size permit/storage on peak-day draw)
Back-end fabric (400G; healthy 1:1 leaf tier)8 rails; 4 × 27-leaf pods/rail; 32 spines/pod864 leaves + 1,024 spines + 512 cores = 2,400 switches; 176,128 module ends
Floor area (white space)384 racks × assumed 30 m²/rack including aisles/CDUs~12,000 m² + plant (two-significant-figure layout allowance)
Floor loading basisdeclared equivalent-uniform screen plus OEM foot/wheel, rolling and rigging reactionsverify project load combinations against the complete slab/access-floor assembly and move route
GPU-rack fixture: 132 kW/rack. The PUE ≈ 1.2 multiplication screens this fixture; add supporting IT before sizing facility demand. The protected-IT transient boundary and generator-terminal critical-facility boundary are separate: rack-side peaks are allocated across capacitance/BBU/BESS/UPS, while standby generation carries the declared steady island load plus explicit derating and reserve. Water assumes hybrid rejection with adiabatic assist; dry, non-evaporative heat rejection trades water against site-specific energy.

Campus port map. In each rail/pod, leaf ports 1–32 face hosts and 33–64 face spines 1–32. Each spine uses 1–27 for leaves, 28–32 dark, 33–48/49–64 for its two group cores. Each core’s four consecutive 16-port groups face the four pods’ matching spine. Every cable end has one address. The ledger totals 27,648 host + 27,648 leaf–spine + 32,768 spine–core = 88,064 channels, 176,128 module ends and 704,512 active fiber paths; surveyed lengths determine fiber-kilometers, and electrical lanes are separate.

At 108 leaves/rail the campus exceeds Chapter 8.5’s 64-leaf two-tier limit. Include the core; at ≤16,384 GPUs its two-tier instance becomes eligible. Release only after Chapter 13.7/Chapter 13.9 pass the purchased stack and failure states. A spare chassis repairs an outage; it cannot keep a single-attached rank connected.

4. Reference BOM: a 100k-GPU GB200 cluster

The flagship build: a 100,000-GPU GB200 reference training cluster, costed as a bill of materials. Built from the SU: ceil(100,000 ÷ 576) = 174 SUs minimum; choose eleven 16-SU halls, hence 176 SUs = 1,408 NVL72 racks = 101,376 GPUs (≈ 100k). At 132 kW/rack that is ~186 MW GPU-rack IT and ~220 MW facility for that fixture at assumed PUE 1.2. Add separately counted supporting IT before a project capacity release. That capacity may be delivered on one qualified campus or split across sites; DCI does not imply that this illustrative logical cluster runs as one cross-site synchronous job (Chapter 8.8).

The figures use SemiAnalysis’s August 2025 server-only estimate and explicitly assumed storage/CDU/plant costs. Use the counts as arithmetic — they balance ports, power, and flow; they do not draw a topology — and the dollars as an order-of-magnitude frame. The GPU/rack line is a major cost; its share requires priced networking and reconciled facility scope. A percentage allowance cannot supply the core.

100k-GPU cluster reference BOM (176 SUs · 1,408 NVL72 racks · 101,376 GPUs)
BOM lineCountBasisUnit-cost basisLine cost (rough)
GB200 GPUs (rack population)101,376176 SU × 576Included in the integrated rack lineDo not add a bundled network allowance to the counted fabric
NVL72 racks (integrated, L11)1,408176 × 8Server estimate belowSubtotal below
— includes GPUs, Grace, NVSwitch and DLC—101,376 × unrounded per-GPU server cost = same subtotalEquivalent server scopeIncluded above
CDUs (HPE 1.3 MW; maximum eight racks/CDU; local 3+1 pods)235 installed1,408 racks: 176 duty; 58 full 24-rack 3+1 pods plus one final 16-rack 2+1 podIllustrative unit allowance below235 × allowance; rounded subtotal below
Back-end leaves (64×400G)3,1688 rails × 396 leaves; 32 host + 32 uplink portsAcquire switch quote3,168 × quote
Back-end spines and cores3,328 + 2,048Per rail: 13 × 32 spines; 32 × 8 cores; 4 links/spine/coreAcquire both quotesPrice both tiers and module power
Front-end/storage; separate OOBAcquire port schedules9.3 service/recovery; 8.7 management inventoryAcquire quotesSeparate fabric scope
Back-end modules (assumed 400G DR4)618,496 ends101,376 host + 101,376 leaf–spine + 106,496 spine–core channels; two ends/channelAcquire module/channel quotes309,248 channels plus spares
DAC/AEC copper (intra-rack scale-up)in-rack5,184 NVLink cables/rack (copper, in rack price)incl. in rack—
Hot storage (parallel FS)~200 PB usable99 × 2 PB = 198 PB usable capacity; bandwidth HOLD for workload budgetsAssumed usable-GB price below198M GB × price; subtotal below
Capacity / object tier~150–250 PBData lake + checkpoints; independent endpointsAssumed usable-GB price below150M × low price to 250M × high price
Facility power chain~220 MW facilityTransformers, UPS/BESS, switchgear, generatorsAssumed facility-MW price belowUnrounded MW × price; excludes shell
Cooling plant + water loop~186 MW ITRejection/piping/plant; excludes 235 CDUsAssumed IT-MW price belowUnrounded MW × price; excludes shell
Cluster totalHOLD for complete scopeReconcile server/facility scopes; include every fabric tier, passive plant, power and sparesAcquire quotesAn unpriced core cannot fund a build
Equipment counts follow the stated reference inputs; networking follows the complete port map below. Unit costs and rounded subtotals below; assumed CDU/media/plant bands explained above. Excludes land, shell and soft costs. Networking is the generic 400G three-tier scenario: eight rails, 64-port switches, healthy 1:1 leaves. All switches/module ends are counted. Dark ports serve partial pods and upper-tier capacity; add cold spares, passive plant and installed power. Storage bandwidth remains HOLD for the independent loader, checkpoint drain/commit and restore budgets in Chapters 9.4/9.8. Compare the GB200 server estimate with Appendix C’s GB300 figures below, a different generation and complete-system boundary, per 72-GPU system and per installed GPU — do not read the two as the same number.

Large-cluster port map. Leaves keep the campus assignment. In each rail, twelve pods use spine ports 1–32 for leaves; the final pod uses 1–12, with 13–32 dark. Spine ports 33–64 face eight group cores, four ports/core. Each core’s thirteen four-port pod groups occupy 1–52; 53–64 stay dark. This maps 101,376 endpoints to 8,544 switches; all channel/module populations are in the BOM. It is a realizable map, not a minimum-switch proof.

A campus core loss removes half its group’s upper capacity; a cluster core loss removes one eighth. A leaf loss disconnects 32 ranks. Test these cuts under Chapter 8.5 with actual placement and payload service; healthy port balance does not accept a failure. Budget release stays HOLD for route, recovery, power and quote evidence. At ≤16,384 GPUs, qualifying the two-tier map can remove the core purchase.

135 kW nominal TDP / up to 155 kW peak
Lenovo GB300 NVL72: nominal rack TDP and workload-dependent peak
Scope & caveats

One named Lenovo product record. These are distinct operating quantities, not a single scalar or a project electrical design rating.

Qualify the purchased configuration and transient envelope. The separate NVIDIA reference architecture figure of 142 kW is not a Lenovo design provision.

576 GPUs / SU
DGX SuperPOD GB200 scalable unit = 8 NVL72 racks; full SuperPOD = 16 SU / 128 racks / 9,216 GPUs
~600 kWforecast
Rubin Ultra Kyber rack (NVL144) on 800 VDC; ~4.8 MW per 8-rack SU (an NVL1152 domain)
Scope & caveats

NVIDIA's published figure (GTC 2025) is 600 kW per Rubin Ultra Kyber rack and GTC 2026 did not revise it. SemiAnalysis (2026-05-26) reports Kyber Ultra 'approaching 660 kW' — a single-source analyst estimate for a 2027 part, recorded here rather than adopted, since the vendor primary figure still stands.

~1.43 L/min/kWderived
guide water heat-balance at 10 K: ~1.43 L/min per kW of liquid-captured heat; project flow uses the approved fluid and selected ΔT
2025Guide heat-balance derivation using water properties at the declared design pointregister ↗
Scope & caveats

A guide physics result, not an OCP or OEM flow specification. Apply it only to liquid-captured heat and recalculate with the approved fluid properties, selected ΔT, pressure budget, and product operating envelope.

1.3 MW
HPE GB200 NVL72: 1.3 MW CDU, maximum eight racks; Appendix B uses 48 duty plus 16 local standbys for 384 racks
8× 400 Gb/s
named 8-rail reference: 8×400 Gb/s per node; validate ports, rails, and blocking ratio against measured traffic and step-time targets
Scope & caveats

One named 8-rail training reference design, not a workload-name mandate. Validate the port rate, rail count, and blocking ratio against the target traffic matrix, topology, placement, failure headroom, and step-time objective.

$3.1M/rack; about $43k/GPUderived
SemiAnalysis GB200 server-only estimate and per-GPU equivalent; price the fabric separately
Scope & caveats

Hyperscaler server only; rack and per-GPU costs equivalent. Network/storage/facility extra.

~$6.25M/system; ~$86.8k/installed GPUderived
GB300 complete-system default: 72 GPUs (Appendix C)
Sep 2026Guide derivation from Appendix C’s assumed GB300 complete-system default: $86,800 per installed GPU × 72 installed GPUs; not an observed market quote.register ↗
Scope & caveats

Assumed GB300 complete system; network/facility separate.

about $4.4Bderived
GB200 server subtotal: 1,408 racks = 101,376 GPUs
Scope & caveats

Server only; add separately quoted network, storage and facility.

$120–180k each; about $28–42Mestimate
235 CDUs: assumed unit price → subtotal
Sep 2026Guide derivation from explicitly assumed unit-price and capacity sensitivities; these are illustrative allowances, not observed quotes.register ↗
Scope & caveats

Guide allowance; count CDUs once (Chapter 5.6).

$0.20–0.40/GB; about $40–80Mestimate
198 PB usable flash: assumed price → subtotal
Sep 2026Guide derivation from explicitly assumed unit-price and capacity sensitivities; these are illustrative allowances, not observed quotes.register ↗
Scope & caveats

Guide allowance, decimal usable GB (Chapters 9.6/9.8).

$0.02–0.05/GB; about $3–10Mestimate
150–250 PB object tier: independent assumed endpoints
Sep 2026Guide derivation from explicitly assumed unit-price and capacity sensitivities; these are illustrative allowances, not observed quotes.register ↗
Scope & caveats

Guide allowance; unrounded high endpoint $12.5M, not a $10M cap (Chapter 9.6).

$10–15M/MW; about $2.2–3.3Bestimate
Facility power chain: assumed price → subtotal
Sep 2026Guide derivation from explicitly assumed unit-price and capacity sensitivities; these are illustrative allowances, not observed quotes.register ↗
Scope & caveats

Guide allowance; exclude shell and cooling (Chapter 4.1).

$3–5M/IT MW; about $0.6–0.9Bestimate
Cooling plant excluding CDUs: assumed price → subtotal
Sep 2026Guide derivation from explicitly assumed unit-price and capacity sensitivities; these are illustrative allowances, not observed quotes.register ↗
Scope & caveats

Guide allowance; exclude the 235 CDUs and shell (Chapter 5.8).

Sensitivity: how the two builds move when you change one input

Generation step (GB200 → Vera Rubin). A future 8-rack SU must be recalculated from that named product's liquid-captured heat, approved-fluid properties and selected ΔT; the current GB200 illustration does not supply a portable flow endpoint for Vera Rubin. A hard 50 MW operating envelope at Vera Rubin density fits 33 full SUs at 188 kW Max Q (1.504 MW/SU) or 27 at 228 kW Max P (1.824 MW/SU); a 50 MW facility-basis envelope fits 18 at the 330 kW cabinet basis (2.64 MW/SU). At the 188 kW Max-Q point, nominal heat per rack is about 1.4× the 132 kW GB200 TDP fixture; the 228 kW Max-P and 330 kW facility bases are separate comparisons. Select the generation only after cooling, electrical capacity and the floor/move route pass together. Oversubscription scenario (1:1 → 2:1). Recount the 1:1 → 2:1 equipment change under 8.5, then test surviving cuts in 13.7 and workload deadlines in 13.9. Accept savings only after placement and failure states pass; another topology’s percentage cannot price this core. Dry, non-evaporative heat rejection. Moves the campus’s illustrative ~200 ML/year evaporative make-up toward zero, with loop fill and maintenance still required. At the assumed +0.05 PUE sensitivity, 50.688 MW × 0.05 adds about 3 MW of facility load; the named weather/plant curve decides that increment and the larger rejection footprint. Reject dry cooling if that load exceeds the POI or lifecycle-cost budget; Chapter 15.4 owns the WUE/PUE trade. Effective GPU life (3 yr → 5 yr). At constant installed population and utilization, 5/3 increases lifetime GPU-hours by about 67%. Annual straight-line depreciation falls to 3/5 of its former value; recurring energy and operating cost do not fall by that factor. No equipment count or initial capex changes. Use Chapter 1.8 and Appendix C to test whether the hardware remains economically useful for those extra years.

Choose the reference population whose counted fabric and surviving workload fit the project; procure the complete scope. 8.5 changes topology; 1.8 values it. An incomplete fabric allowance leaves required hardware, power and commissioning outside the purchase. Close the supporting-IT inventory with 7.13 and the counted fabric in 8.5; then 4.1 owns electrical demand and surviving source capacity, while 5.11 owns cooling flow and heat rejection with a train unavailable. The integrated release needs the actual storage, management and network load, qualified failure state and non-overlapping quotes; rack nominal MW alone releases none of those interfaces.

These reference designs operationalize the archetype framework in Chapter 1.1 and the requirements matrix in Chapter 1.7; the economics that score them live in Chapter 1.8 with the calculators in Appendix C. The SU and BOM inherit: rack/integration detail from Chapter 7.13 and Chapter 7.14; the 800 VDC power chain from Chapter 4.7 and transient sizing from Chapter 4.5; CDU and warm-water loop sizing from Chapter 5.6 and Chapter 5.7; fabric topology and oversubscription from Chapter 8.5 and optics from Chapter 8.10; storage sizing from Chapter 9.8; multi-campus scale-across from Chapter 8.8. Every dated figure here is registered with vintage and scenario in Appendix D.
Cite this chapter
Fehn, J. (2026). Reference Designs & Worked Examples (Chapter B). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/appendix-appendices-and-reference-data/b-reference-designs-and-worked-examples (accessed 2026-09-29).
@misc{aidc-B,
  author       = {Fehn, Jacob},
  title        = {Reference Designs & Worked Examples (Chapter B)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/appendix-appendices-and-reference-data/b-reference-designs-and-worked-examples},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit