GPU · Graphics Processing Unit
A massively parallel processor that became the workhorse of AI training and inference.
Current numbers
186 GBGB200 usable HBM per GPU
~1–3% tuned servingBlackwell (B200) confidential-computing overhead for a tuned LLM serving stack — 30–40% for a stock stack, ~10–13% for eight-GPU CC training
under 45 min (4 GPU) / under 2.25 hr (8 GPU)DCGM level-4 estimated runtime, current documentation (4-GPU vs 8-GPU systems) — measure it on your own node type