The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

share of Meta Llama 3 unexpected training interruptions from faulty GPUs / HBM3 memory (148 and 72 of 419)

35.3% / 17.2%derived

Value kindderived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method.
ScopeOne 54-day pre-training window on a 16,384-GPU H100 cluster — a single fleet, generation, and workload, not an industry rate.
CaveatGuide arithmetic on the paper's own counts: 148/419 = 35.3% and 72/419 = 17.2%. Table 5 prints 30.1% for the faulty-GPU row, which no denominator in the paper reproduces; its 18 rows do sum to 419, but the printed percentages sum to 94.9%. Counts are auditable; treat the printed percentages as paper-printed.
As of2024-07
SourceMeta AI, The Llama 3 Herd of Models, arXiv v1, July 31, 2024, §3.3.4 / Table 5. · §3.3.4 and Table 5: 466 total interruptions, 47 planned, 419 unexpected; Faulty GPU 148, GPU HBM3 Memory 72, Software Bug 54, Network Switch/Cable 35
Reviewchecking…review by 2026-12-22 · standard cadence
Recorded changeslast 2026-09-16 · 5 revisions tracked
Claim idshare-of-training-interruptions-from-faulty

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.