The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

projected mean time to failure for a 16,384-GPU synchronous job, population-scaled from Meta research-cluster data (~7.9 hr observed at 1,024 GPUs)

~1.8 hr (projection)forecast

Value kindforecast — Forecasts are time-bound outlooks, not specifications; test decisions across plausible scenarios.
ScopeA projection from RSC-1/RSC-2 job-failure populations, not an observation. The separately observed Llama 3 405B run on 16,384 H100s (419 interruptions in 54 days) averaged ~3 hr between interruptions with a different event population; neither extrapolates to a 131,072-GPU forecast without explicit job-membership and common-mode assumptions.
As of2024
SourceMeta, Revisiting Reliability in Large-Scale ML Research Clusters
Reviewchecking…review by 2026-08-17 · standard cadence
Recorded changeslast 2026-09-01 · 3 revisions tracked
Claim idmean-time-to-failure-for-a-16-384-gpu

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.