The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuideGlossaryDecode

Decode

The memory-bandwidth-bound phase that generates output tokens one at a time after prefill.

Current numbers

GTC 2026 (reported)Groq 3 LPU chip (~500 MB SRAM, ~150 TB/s) → LPX rack (256 LPUs, ~128 GB aggregate SRAM) fills the disaggregated decode/FFN slot (no HBM); Rubin GPUs retain prefill + attention. Rubin CPX reportedly pulled at GTC 2026 (no official NVIDIA cancellation)as of 2026-07 · register ↗
xPyDruntime-reconfigurable disaggregation: x prefill workers feeding y decode workers, re-balanced liveas of 2026 · register ↗
10–100xper-sample CPU decode cost of image/video vs text — the reason the GPU:CPU ratio must be set per modalityas of 2025 · register ↗

← All terms