MLPerf
The industry-standard benchmark suite for AI training and inference performance; the closest thing to apples-to-apples numbers across accelerator vendors.
Current numbers
≥90%accelerator-utilization threshold for a passing MLPerf Storage result (70% for Cosmoflow); holding this line is the loader's core objective
~every 30 minMLPerf Storage model for 100k accelerators at full utilization: ~one failure per 30 min; illustrative checkpoint workload, not an observed fleet rate