mean-time-to-failure of a 1,024-GPU job vs 47.7 days for an 8-GPU job — the single-point-of-failure penalty of scale
7.9 hrobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| As of | 2025 |
| Source | Meta, Revisiting Reliability in Large-Scale ML Clusters (arXiv 2410.21680) |
| Review | checking…review by 2026-08-18 · standard cadence |
| Recorded changes | last 2026-06-29 |
| Claim id | mean-time-to-failure-of-a-1-024-gpu-job-vs-47-7 |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.