MI300X
AMD's first broadly deployed AI GPU and the canonical case study in the gap between paper FLOPS and realized throughput once the software stack is accounted for.
Current numbers
~20%MI300X cost-per-token advantage vs H100 SXM in small-batch / largest-model inference regimes
192 GBMI300X HBM3 @ 5.3 TB/s, 750 W (CDNA 3) — 2.4x H100 capacity; MI325X 256 GB @ 6 TB/s, 1,000 W