MMLU · Massive Multitask Language Understanding
A broad knowledge benchmark used as the quality yardstick when judging whether a lower-precision training or quantization recipe has degraded a model.
Current numbers
25T tokenslongest documented 4-bit pre-training run (12B Mamba-Transformer, NVFP4); 62.58% vs 62.62% MMLU-Pro vs FP8
≤1%NVFP4 inference accuracy drop vs FP8 on DeepSeek-R1 (MMLU-Pro 84% vs 85%)
~29%estimated MMLU contamination across public web corpora; clean-mirror retests drop scores high-single to low-double digits