Model FLOPs Utilization · MFU
How much of a chip's theoretical compute a training run actually uses; the headline training-efficiency metric.
Current numbers
~15–30%AMD Instinct hardware price discount to NVIDIA SXM — real only if the realized-MFU gap on your workload is smaller
37–66%MI300X realized inference throughput vs H100/H200 (2025), despite ~1.3x paper FLOPS — the historical realized-MFU gap
35–55%typical training MFU band for a well-tuned dense transformer on a mature stack; immature stacks roughly halve it