MFU · Model FLOPs Utilization
Achieved FLOPs divided by peak FLOPs in a training run; 35-55% is good at scale, eroded by collectives and stragglers.
How much of a chip's theoretical compute a training run actually uses; the headline training-efficiency metric.
Also written as: Model FLOPs Utilization
Current numbers
~41%BF16 MFU achieved pre-training Llama 3 on 16k H100s (frontier-scale reference point)
55.2% MFU (named MegaScale run)training MFU on a named MegaScale run (ByteDance / Peking University) — a stricter metric than the device-utilization figures usually quoted for enterprise clusters
34% → 54%BF16 MFU gain on H100 (GPT-3 175B) from software and kernel maturation over ~12 months, Jan–Dec 2024 — same silicon, more throughput