The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Calculators

Open models for the numbers that decide an AI build — one project, worked end to end: size the cluster, price the build, cost the runtime, then see if the finance closes. Defaults trace to the numbers register (current as of 2026-07); change them to your case, then save, share a permalink, or export to CSV.

One project scenario — every calculator derives from it
Workload
Hardware generation
1200 W · 186 GB HBM · 72/rack · 132 kW/rack · 2.5 PFLOPS dense
Facility & location
Financing & schedule
NVIDIA GB200 NVL72208 GPUs · 3 racks132 kW/rack → Direct-to-chip liquid (OEM high-density)0.4 MW IT / 0.4 MW facility$16M program$3.22/GPU-hr · $2.55/M tok1,345 t CO₂/yr95.9% equity IRR · 5.07× min DSCR
All assumptions are the guide defaults (asof 2026-07). Everything runs in your browser — the share link carries the scenario in the URL fragment (never sent to a server); only an explicit Save stores it to your account, versioned (v1).

Size it

Cluster sizing — model to megawatts

208 GPUs · 0.4 MW
VRAM per replica (weights 70 + KV 676 GB)858 GB
GPUs per replica → replicas for demand (VRAM needs 5; TP rounds up to a power of two)8 × 26
Racks (72/rack)3
Facility draw (incl. host + PUE)0.44 MW
Fleet rental cost$2,850,037/mo

Weights are the floor — at long context the KV cache dominates VRAM (70–90% at 1M tokens → Ch 9.7). Tensor-parallel degree prefers powers of two; serving engineering (batching, disaggregation) moves throughput ~10× → Ch 10.11.

Density → cooling regime

Direct-to-chip liquid (OEM-qualified high-density)

D2C cold-plate platforms serve 200–600 kW racks (NVL72 → Kyber-class) with sidecar power and liquid busbars. Density alone never forces immersion — it is an architecture choice for specific niches. Bands are approximate (air ≲30 kW · rear-door bridges to ~100 · direct-to-chip beyond; immersion is a niche architecture, not a density mandate) → Ch 5.4.

Build it

Build cost — capex per megawatt

$0.01B excl. servers
Facility (shell, electrical, cooling, land)$4M
Network & cluster infrastructure$2M
Intensity, servers excluded$16.7/W

Benchmarks: JLL global-average shell/core ≈ $11.3M/MW (50 MW single-tenant, air-cooled; liquid +10%); Epoch's 1 GW model runs facility+land+utility ~$11.8M/MW and network ~$4.9M/MW, with servers ~$21.2M/MW on top → ~$38B/GW up-front. Electrical dominates the facility split → Ch 2.5, Ch 1.8.

Scope & how the benchmarks reconcile

This models construction-period capex excluding the server fleet. The benchmarks, reconciled on Epoch AI's May-2026 1 GW model: $11.3M/MW is JLL's global-average shell-and-core (air-cooled, 50 MW single-tenant; liquid cooling adds ~10%); Epoch's facility + land + utility works line is ~$11.8M/MW and its network and cluster infrastructure ~$4.9M/MW; servers — GPUs included — add ~$21.2M/MW, which is how a 1 GW campus reaches ~$38B up-front. Electrical systems are 45–70% of construction cost (as of 2026 → register). The server fleet is deliberately outside this calculator — the TCO model owns it, and adding it here would double-count.

Site-scoring playbook

64 / 100 · Workable with mitigation

Kill gates come first: power or interconnect at ≤2, or water at ≤1, disqualifies a site no matter how the weighted score averages out. Weights reflect the reordered 2026 hierarchy — power availability now leads, ahead of latency and land. Tune to your workload (training tolerates latency; inference doesn't) → Ch 3.13.

Redundancy topology

Tier 3 · N+1, concurrently maintainable
Concurrently maintainableYes
Fault tolerantNo
Relative capex≈1.4×
Typical useEnterprise / most AI training

Tier certifies which maintenance and fault events the topology rides through — not an availability percentage (Uptime disavowed the legacy %-uptime mappings). Model service availability from your own event/repair data; the capex multiple is indicative (2N roughly doubles electrical/mechanical) → Ch 12.1.

Run it

Training run — time, cost, energy & CO₂

8.1 days · $36.5M rental
Compute (6·P·T)6.3×10²⁴ FLOPs
GPU-hours1.94M
Facility energy (incl. host + PUE)4,100 MWh (~$0.33M)
Emissions at 350 gCO₂/kWh1,435 t CO₂

MFU and goodput compound: 40% MFU × 90% goodput means ~2/3 of paid FLOPs never reach the loss curve. The gap between ~90% and ~96% goodput is checkpointing and cordon policy → Ch 12.2.

GPU TCO & cost-per-GPU-hour

$3.22 / GPU-hr
Annual cost / GPU$19,722
· Depreciation$14,713
· Energy$1,478
· Opex$3,531
Breakeven utilization vs rental12%

Ownership beats the rental rate only above the breakeven utilization computed from your own inputs — below it a debt-financed cluster bleeds cash → Ch 1.8.

Inference cost per million tokens

$14.9 / M tokens

Self-hosted worked example. Market self-serve fell ~$10 → ~$2.50 / M tokens in a year (~4×) — and the cost to serve a fixed-quality token falls faster still (LLMflation ~10×/yr) — so underwrite inference with a price-decline curve, never flat → Ch 1.8.

Facility energy & water

Facility draw0.4 MW
Annual energy3,842 MWh
Annual energy cost$307,324
Annual water1.3 M L

PUE bands: legacy air 1.4–1.6 · direct-to-chip liquid 1.05–1.15 → Ch 15.1.

Fund it

Project finance — CFADS, DSCR, IRR

95.9% equity IRR
Project (unlevered) IRR71.2%
Min / avg DSCR (CFADS ÷ debt service)5.07× / 9.69×
Debt at COD (incl. $1M IDC)$10M
Equity invested (incl. $1M DSRA)$7M
Project NPV @ 10%$143M
Equity multiple47.06×

A real CFADS pro-forma: capex draws over the build with capitalized IDC, a revenue ramp, cash taxes net of the straight-line depreciation and interest shields, and sustaining capex. DSCR is CFADS ÷ debt service — the lender's ratio, stricter than EBITDA coverage; screen against ~1.3–1.4× with margin. Contracted offtake supports more leverage than merchant. Working-capital swings and NOL carryforwards are not modeled → Ch 2.5.

Then turn the sizing into dates: the lead-time planner reverse-schedules every long-lead PO from your ready-for-service target.