MMLU · Massive Multitask Language Understanding
A broad knowledge benchmark used as the quality yardstick when judging whether a lower-precision training or quantization recipe has degraded a model.
Current numbers
1 percentage pointone-percentage-point MMLU-Pro decrease in NVIDIA’s rounded DeepSeek-R1-0528 comparison
~29%estimated MMLU contamination across public web corpora; clean-mirror retests drop scores high-single to low-double digits