MI300X
AMD's first broadly deployed AI GPU and the canonical case study in the gap between paper FLOPS and realized throughput once the software stack is accounted for.
Current numbers
37–66%MI300X realized inference performance vs H100/H200 despite ~1.3x paper FLOPS — the ROCm goodput gap
~20%MI300X cost-per-token advantage vs H100 SXM in small-batch / largest-model inference regimes
37–66%MI300X realized inference throughput vs H100/H200 (2025), despite ~1.3x paper FLOPS — the historical realized-MFU gap