NVFP4
NVIDIA's 4-bit microscaling format, using small value blocks with a floating-point scale; the choice between it and MXFP4 is baked into silicon and into checkpoint portability.
Current numbers
25T tokenslongest documented 4-bit pre-training run (12B Mamba-Transformer, NVFP4); 62.58% vs 62.62% MMLU-Pro vs FP8
~50 PFLOPSVera Rubin per-GPU FP4 inference (Rubin Ultra ~100 PFLOPS NVFP4)
88%per-block MSE reduction from NVFP4's E4M3 scale vs MXFP4's E8M0 (0.72→0.08)