projected mean time to failure for a 16,384-GPU synchronous job, population-scaled from Meta research-cluster data (~7.9 hr observed at 1,024 GPUs)
~1.8 hr (projection)forecast
| Value kind | forecast — Forecasts are time-bound outlooks, not specifications; test decisions across plausible scenarios. |
|---|---|
| Scope | A projection from RSC-1/RSC-2 job-failure populations, not an observation. The separately observed Llama 3 405B run on 16,384 H100s (419 interruptions in 54 days) averaged ~3 hr between interruptions with a different event population; neither extrapolates to a 131,072-GPU forecast without explicit job-membership and common-mode assumptions. |
| As of | 2024 |
| Source | Meta, Revisiting Reliability in Large-Scale ML Research Clusters |
| Review | checking…review by 2026-08-17 · standard cadence |
| Recorded changes | last 2026-09-01 · 3 revisions tracked |
| Claim id | mean-time-to-failure-for-a-16-384-gpu |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.