Checkpoint
A periodic save of model weights and optimizer state so a long training run can resume after a failure.
Current numbers
~1 failure / 512 GPUs / weekbest-in-class H100 cluster failure rate; one failure restarts a synchronous job from checkpoint
~4 GB/s/GPUNVIDIA reference read target for vision training; ~1 GB/s/GPU practical floor; 4–10 GB/s/GPU for real multimodal/checkpoint-heavy runs
~14 bytes/paramcheckpoint size incl. optimizer state — 100B ~1.4 TB; GPT-3 175B ~2.45 TB; 1T params ~14 TB; frontier runs sustain on < 1 TB/s global checkpoint BW