The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Glossary › Decode

Decode

The memory-bandwidth-bound phase that generates output tokens one at a time after prefill.

In other contexts, Decode also means In the data pipeline: turning stored, compressed samples (images, video, packed text shards) back into training-ready tensors — CPU-bound work that gates GPU utilization. (Parts 9.5).

Current numbers

70 billion parameters × 2 B; 80 layers; 8 KV heads; head dimension 128; 2 B/KV value; 8,192 cached tokens; batch 8; workspace 20 GB; usable HBM 186 GB; sustained HBM 4.0 TB/s; step budget 50 ms.Fit one decode replica before buying more HBM — input ledgeras of 2026-09 · register ↗
About 180 GB resident; 40 ms memory lower bound. At most 9 sequences fit; 10 fail. Below about 3.2 TB/s the 50 ms memory budget fails.Fit one decode replica before buying more HBM — result and flip thresholdas of 2026-09 · register ↗
GTC 2026 (reported)Groq 3 LPU chip (~500 MB SRAM, ~150 TB/s) → LPX rack (256 LPUs, ~128 GB aggregate SRAM) fills the disaggregated decode/FFN slot (no HBM); Rubin GPUs retain prefill + attention. Rubin CPX reportedly pulled at GTC 2026 (no official NVIDIA cancellation)as of 2026-07 · register ↗

← All terms