failure cadence in a 100k-accelerator cluster at full utilization — why checkpoint bandwidth is an acceptance gate
~every 30 minderived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | MLCommons MLPerf Storage v2.0 illustrative checkpoint benchmark model for 100,000 accelerators at full utilization, not an observed fleet failure rate. Size a real cluster's checkpoint interval from its own measured interruption distribution. |
| Caveat | Illustrative benchmark workload, not a fleet forecast. Common-mode events, software, repair, job membership, censoring, detection policy, and non-constant hazard are outside this calculation. |
| As of | 2025 |
| Source | MLCommons MLPerf Storage v2.0 · Section “Version 2.0 adds checkpointing tasks”: a 100,000-accelerator cluster with 50,000-hour accelerator MTTF will likely experience a failure every half-hour. |
| Derivation | 100,000 / 50,000 accelerator-hours = 2 expected accelerator failures per hour under the stated independent constant-rate model |
| Review | checking…review by 2026-11-23 · standard cadence |
| Recorded changes | last 2026-06-29 |
| Claim id | failure-cadence-in-a-100k-accelerator-cluster |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.