NVFP4
NVIDIA's 4-bit microscaling format, using small value blocks with a floating-point scale; the choice between it and MXFP4 is baked into silicon and into checkpoint portability.
Current numbers
25T tokenslongest documented 4-bit pre-training run (Nemotron 3 Super, 120B-total / ~12B-active hybrid Mamba-Transformer MoE, NVFP4, 25T total seen tokens)
~50 PFLOPSVera Rubin per-GPU FP4 inference (Rubin Ultra ~100 PFLOPS NVFP4)
16 vs 32NVFP4 block size (16 values, FP8 scale) vs MXFP4 (32 values, power-of-two scale)