Fit one decode replica before buying more HBM — stated inputs
70 billion parameters × 2 B; 80 layers; 8 KV heads; head dimension 128; 2 B/KV value; 8,192 cached tokens; batch 8; workspace 20 GB; usable HBM 186 GB; sustained HBM 4.0 TB/s; step budget 50 ms.modeled
| Value kind | modeled — Interpret this value according to its displayed kind, scope, source, and as-of date. |
|---|---|
| Scope | Synthetic GQA shape; decimal GB/TB, no offload or sharing. NVIDIA supplies usable capacity; other workload inputs and sensitivities are unsupported teaching assumptions explained in the opening callout. Chapter 1.7 owns workload definition. |
| As of | 2026-09 |
| Source | NVIDIA GB200 specifications and CUDA performance memory model — method reference; guide-authored hypothetical scenario September 8, 2026. No supplier quote or test measurement. |
| Derivation | Exact teaching assumptions chosen to expose the stated constraint and its reversal; not estimates of market prices or actual equipment. |
| Review | checking…review by 2027-03-08 · standard cadence |
| Recorded changes | last 2026-09-16 |
| Claim id | guide2-t29-7-6-inputs |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.