Illustrative mixed-precision Adam persistent state with BF16 weights/gradients, FP32 master weights and two FP32 moments
16 B/parameter for the stated layoutderived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | 2 B weights + 2 B gradients + 4 B FP32 master weights + 8 B Adam moments; excludes activations, temporary gathers, buffers and allocator headroom. Not a universal framework default. |
| As of | 2026-09-08 |
| Source | Guide, persistent-state representation arithmetic, 8 September 2026; DeepSpeed memory documentation supplies the representation context. · Memory Requirements discussion: FP16 parameters use 2 bytes, FP16 gradients 2 bytes, Adam momentum and variance 8 bytes total, with FP32 master parameters bringing the standard total to approximately 16 bytes per parameter |
| Derivation | 2 + 2 + 4 + 2 * 4 = 16 B/parameter. |
| Review | checking…review by 2027-03-05 · standard cadence |
| Recorded changes | last 2026-09-16 · 2 revisions tracked |
| Claim id | training-state-footprint-with-mixed-precision |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.