Cluster sizing — model → MW
Size an inference fleet end to end: model parameters, quantization, KV cache and context length → VRAM per replica → active GPUs → installed racks, stranded capacity, facility megawatts, and fractional rental cost.
Cluster sizing — model to megawatts
VRAM establishes only a memory floor. Enter a replica degree qualified for the named model, runtime, hardware, context and batch profile; supported degrees need not be powers of two. Purchase quantum is a separate SKU or contract input and is not inferred from rack density. Facility design power follows installed units, while rental cost follows the active fractional fleet. At long context the KV cache can dominate VRAM → Ch 9.7.
These are transparent screening estimates, not final designs, financial advice, or project approvals. Replace every default with current project data and have the responsible project authorities validate the result. You can save, share a permalink, or export to CSV; see the full calculator suite.