The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuideGlossaryCheckpoint

Checkpoint

A periodic save of model weights and optimizer state so a long training run can resume after a failure.

Current numbers

~1 failure / 512 GPUs / weekbest-in-class H100 cluster failure rate; one failure restarts a synchronous job from checkpointas of 2025 · register ↗
~4 GB/s/GPUNVIDIA reference read target for vision training; ~1 GB/s/GPU practical floor; 4–10 GB/s/GPU for real multimodal/checkpoint-heavy runsas of 2025 · register ↗
~14 bytes/paramcheckpoint size incl. optimizer state — 100B ~1.4 TB; GPT-3 175B ~2.45 TB; 1T params ~14 TB; frontier runs sustain on < 1 TB/s global checkpoint BWas of 2025 · register ↗

← All terms