The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

assumed serialized-state recipe: 2B BF16 weights + 4B FP32 master + 4B+4B Adam moments = 14 bytes/param; ~7x the weight file

~14 bytes/paramderived

Value kindderived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method.
ScopeAn arithmetic recipe for the standard BF16-compute / FP32-master / Adam mixed-precision configuration, not a measured population value. VAST Data assumed 14 bytes/param to infer model sizes from checkpoint sizes (its own footnote puts the resulting uncertainty at about ±15%); the survey did not establish the recipe. 8-bit Adam moments give 8 bytes/param; optimizers that drop the second moment change it again. Verify the tensors the chosen framework actually serializes.
As of2026
SourceVAST Data, 'Optimizing Checkpoint Bandwidth for LLM Training' — states 14 bytes/parameter as an assumption used to estimate parameter counts (footnote 1); the recipe itself is arithmetic
Reviewchecking…review by 2027-09-05 · standard cadence
Recorded changeslast 2026-09-04 · 2 revisions tracked
Claim idcheckpoint-state-per-parameter-2b-bf16-weights

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.