longest documented NVFP4 pre-training run — Nemotron 3 Super, a 120B-total / ~12B-active hybrid Mamba-Transformer MoE, 25T total seen tokens
25T tokensobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Caveat | Nemotron 3 Super: 120B total / ~12B active hybrid Mamba-Transformer MoE, 25T total seen tokens pretrained in NVFP4. This is a DIFFERENT model from the earlier dense 12B run, so the 62.58% vs 62.62% MMLU-Pro comparison does not carry over and has been dropped. |
| As of | 2026-03 |
| Source | NVIDIA · Abstract and results tables; 12B model trained on 10 trillion tokens, with 62.58% NVFP4 versus 62.62% FP8 MMLU-Pro |
| Review | checking…review by 2026-11-23 · standard cadence |
| Recorded changes | last 2026-07-27 · 2 revisions tracked |
| Claim id | longest-documented-4-bit-pre-training-run-12b |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.