The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
CalculatorsCluster sizing — model → MW

Cluster sizing — model → MW

Size an inference fleet end to end: model parameters, quantization, KV cache and context length → VRAM per replica → active GPUs → installed racks, stranded capacity, facility megawatts, and fractional rental cost.

ScenarioNVIDIA GB200 NVL72 · 56 active / 72 installed GPUs · 0.13 MW IT · $5M programguide defaults (asof 2026-07)Edit the shared scenario →

Cluster sizing — model to megawatts

56 active / 72 installed GPUs · 0.15 MW design
VRAM per replica (weights 70 + KV 676 GB)858 GB
Qualified GPUs per replica → replicas for demand (VRAM floor 5)8 × 7
Installed GPUs / stranded slots72 / 16
Installed racks (72/rack)1
Facility design draw on installed units (incl. host + PUE)0.15 MW
Rental cost on active GPUs$767,318/mo

VRAM establishes only a memory floor. Enter a replica degree qualified for the named model, runtime, hardware, context and batch profile; supported degrees need not be powers of two. Purchase quantum is a separate SKU or contract input and is not inferred from rack density. Facility design power follows installed units, while rental cost follows the active fractional fleet. At long context the KV cache can dominate VRAM → Ch 9.7.

These are transparent screening estimates, not final designs, financial advice, or project approvals. Replace every default with current project data and have the responsible project authorities validate the result. You can save, share a permalink, or export to CSV; see the full calculator suite.