MegaScale (ByteDance and Peking University) measured 55.2% model-FLOPs utilization for a named 175B-parameter training run on 12,288 GPUs; utilization claims must name metric, workload, window and denominator
55.2% MFU (named MegaScale run)observed
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Do not compare MFU, SM activity, tensor activity, allocation utilization and goodput as one metric. |
| As of | 2024-02 |
| Source | ByteDance and Peking University, MegaScale (arXiv 2402.15627) · Reported 55.2% MFU for a 175B model trained on 12,288 GPUs |
| Review | checking…review by 2026-09-23 · fast cadence |
| Recorded changes | last 2026-09-04 · 4 revisions tracked |
| Claim id | gpu-utilization-in-a-well-orchestrated-ai |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.