The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

projected mean time to failure (MTTF) for a 16,384-GPU vs 131,072-GPU synchronous job, fitted to Meta's A100 research clusters

1.8 hr → 14 minmodeled

Value kindmodeled — Interpret this value according to its displayed kind, scope, source, and as-of date.
ScopeProjection from a failure model fitted to 11 months of two A100 research clusters (RSC-1/RSC-2, >150M A100 GPU-hours). The paper projects 1.8 hours at 16,384 GPUs and 0.23 hours at 131,072 GPUs. Not a measurement, not an H100/Blackwell-era observation, and distinct from Meta's separately published 16,384-H100 Llama 3 run.
As of2024
SourceMeta A100 research-cluster model. This is a hypothetical scale result, not observed H100 behavior or a dated deployment forecast. · Figure 7 and §III: fitted MTTF projection by job size
Reviewchecking…review by 2027-05-01 · standard cadence
Recorded changeslast 2026-09-16 · 2 revisions tracked
Claim idprojected-mean-time-between-failures-for-a-16

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.