The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

GPU- and HBM-attributed interruptions in one Llama 3 405B run: 148 GPU and 72 HBM events among 419 unplanned interruptions over 54 days on 16,384 H100s

148 GPU + 72 HBM interruptionsobserved

Value kindobserved — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range.
ScopeJob-interruption event counts for one named run. The source does not establish unique failed devices or equipment population-time exposure, so these counts must not be annualized into component AFR, fleet lambda, cumulative equipment risk or spares demand.
As of2024
SourceMeta (The Llama 3 Herd of Models, arXiv 2407.21783) · Table 5, reliability analysis: 419 unplanned interruptions over 54 days; 148 faulty-GPU and 72 HBM3-attributed events
Reviewchecking…review by 2026-09-26 · fast cadence
Recorded changeslast 2026-08-24 · 2 revisions tracked
Claim idcombined-h100-gpu-hbm-annualized-failure-rate-1

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.