NVFP4 packed tensor payload plus block-scale accounting relative to 16-bit and 8-bit unscaled payload
~3.6× / ~1.8× payload ratioderived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | Payload and block scales, not total model/service memory; excludes runtime, KV, mixed-precision tensors and layout padding. |
| As of | 2025-06 |
| Source | NVIDIA — Memory discussion: 4.5 bits/value plus FP32 per-tensor scale. · Memory discussion: 4.5 bits/value plus FP32 per-tensor scale |
| Derivation | 16/4.5=3.5556; 8/4.5=1.7778, ignoring per-tensor overhead and assuming packed storage. |
| Review | checking…review by 2027-06-24 · standard cadence |
| Recorded changes | last 2026-09-16 · 2 revisions tracked |
| Claim id | nvfp4-memory-reduction-vs-fp16-vs-fp8 |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.