MFU · Model FLOPs Utilization
Achieved FLOPs divided by peak FLOPs in a training run; 35-55% is good at scale, eroded by collectives and stragglers.
Current numbers
~15–30%AMD Instinct hardware price discount to NVIDIA SXM — real only if the realized-MFU gap on your workload is smaller
37–66%MI300X realized inference throughput vs H100/H200 (2025), despite ~1.3x paper FLOPS — the historical realized-MFU gap
35–55%typical training MFU band for a well-tuned dense transformer on a mature stack; immature stacks roughly halve it