Calculators
Open models for the numbers that decide an AI build — one project, worked end to end: size the cluster, price the build, cost the runtime, then see if the finance closes. Defaults trace to the numbers register (current as of 2026-07); change them to your case, then save, share a permalink, or export to CSV.
Size it
Cluster sizing — model to megawatts
Weights are the floor — at long context the KV cache dominates VRAM (70–90% at 1M tokens → Ch 9.7). Tensor-parallel degree prefers powers of two; serving engineering (batching, disaggregation) moves throughput ~10× → Ch 10.11.
Density → cooling regime
D2C cold-plate platforms serve 200–600 kW racks (NVL72 → Kyber-class) with sidecar power and liquid busbars. Density alone never forces immersion — it is an architecture choice for specific niches. Bands are approximate (air ≲30 kW · rear-door bridges to ~100 · direct-to-chip beyond; immersion is a niche architecture, not a density mandate) → Ch 5.4.
Build it
Build cost — capex per megawatt
Benchmarks: JLL global-average shell/core ≈ $11.3M/MW (50 MW single-tenant, air-cooled; liquid +10%); Epoch's 1 GW model runs facility+land+utility ~$11.8M/MW and network ~$4.9M/MW, with servers ~$21.2M/MW on top → ~$38B/GW up-front. Electrical dominates the facility split → Ch 2.5, Ch 1.8.
Scope & how the benchmarks reconcile
This models construction-period capex excluding the server fleet. The benchmarks, reconciled on Epoch AI's May-2026 1 GW model: $11.3M/MW is JLL's global-average shell-and-core (air-cooled, 50 MW single-tenant; liquid cooling adds ~10%); Epoch's facility + land + utility works line is ~$11.8M/MW and its network and cluster infrastructure ~$4.9M/MW; servers — GPUs included — add ~$21.2M/MW, which is how a 1 GW campus reaches ~$38B up-front. Electrical systems are 45–70% of construction cost (as of 2026 → register). The server fleet is deliberately outside this calculator — the TCO model owns it, and adding it here would double-count.
Site-scoring playbook
Kill gates come first: power or interconnect at ≤2, or water at ≤1, disqualifies a site no matter how the weighted score averages out. Weights reflect the reordered 2026 hierarchy — power availability now leads, ahead of latency and land. Tune to your workload (training tolerates latency; inference doesn't) → Ch 3.13.
Redundancy topology
Tier certifies which maintenance and fault events the topology rides through — not an availability percentage (Uptime disavowed the legacy %-uptime mappings). Model service availability from your own event/repair data; the capex multiple is indicative (2N roughly doubles electrical/mechanical) → Ch 12.1.
Run it
Training run — time, cost, energy & CO₂
MFU and goodput compound: 40% MFU × 90% goodput means ~2/3 of paid FLOPs never reach the loss curve. The gap between ~90% and ~96% goodput is checkpointing and cordon policy → Ch 12.2.
GPU TCO & cost-per-GPU-hour
Ownership beats the rental rate only above the breakeven utilization computed from your own inputs — below it a debt-financed cluster bleeds cash → Ch 1.8.
Inference cost per million tokens
Self-hosted worked example. Market self-serve fell ~$10 → ~$2.50 / M tokens in a year (~4×) — and the cost to serve a fixed-quality token falls faster still (LLMflation ~10×/yr) — so underwrite inference with a price-decline curve, never flat → Ch 1.8.
Facility energy & water
PUE bands: legacy air 1.4–1.6 · direct-to-chip liquid 1.05–1.15 → Ch 15.1.
Fund it
Project finance — CFADS, DSCR, IRR
A real CFADS pro-forma: capex draws over the build with capitalized IDC, a revenue ramp, cash taxes net of the straight-line depreciation and interest shields, and sustaining capex. DSCR is CFADS ÷ debt service — the lender's ratio, stricter than EBITDA coverage; screen against ~1.3–1.4× with margin. Contracted offtake supports more leverage than merchant. Working-capital swings and NOL carryforwards are not modeled → Ch 2.5.
Then turn the sizing into dates: the lead-time planner reverse-schedules every long-lead PO from your ready-for-service target.