projected mean time to failure (MTTF) for a 16,384-GPU vs 131,072-GPU synchronous job, fitted to Meta's A100 research clusters
1.8 hr → 14 minmodeled
| Value kind | modeled — Interpret this value according to its displayed kind, scope, source, and as-of date. |
|---|---|
| Scope | Projection from a failure model fitted to 11 months of two A100 research clusters (RSC-1/RSC-2, >150M A100 GPU-hours). The paper projects 1.8 hours at 16,384 GPUs and 0.23 hours at 131,072 GPUs. Not a measurement, not an H100/Blackwell-era observation, and distinct from Meta's separately published 16,384-H100 Llama 3 run. |
| As of | 2024 |
| Source | Meta A100 research-cluster model. This is a hypothetical scale result, not observed H100 behavior or a dated deployment forecast. · Figure 7 and §III: fitted MTTF projection by job size |
| Review | checking…review by 2027-05-01 · standard cadence |
| Recorded changes | last 2026-09-16 · 2 revisions tracked |
| Claim id | projected-mean-time-between-failures-for-a-16 |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.