Prefill
The compute-heavy phase that processes an inference prompt and builds its KV cache before generation begins.
Current numbers
GTC 2026 (reported)Groq 3 LPU chip (~500 MB SRAM, ~150 TB/s) → LPX rack (256 LPUs, ~128 GB aggregate SRAM) fills the disaggregated decode/FFN slot (no HBM); Rubin GPUs retain prefill + attention. Rubin CPX reportedly pulled at GTC 2026 (no official NVIDIA cancellation)
~7x / ~10xDynamo + wide-EP MoE throughput on GB200 NVL72 vs B200 (Dynamo 1.0 GA at GTC 2026); NIXL+GPUDirect Storage prefill speedup for long context
xPyDruntime-reconfigurable disaggregation: x prefill workers feeding y decode workers, re-balanced live