share of Meta Llama 3 unexpected training interruptions from faulty GPUs / HBM3 memory (148 and 72 of 419)
35.3% / 17.2%derived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | One 54-day pre-training window on a 16,384-GPU H100 cluster — a single fleet, generation, and workload, not an industry rate. |
| Caveat | Guide arithmetic on the paper's own counts: 148/419 = 35.3% and 72/419 = 17.2%. Table 5 prints 30.1% for the faulty-GPU row, which no denominator in the paper reproduces; its 18 rows do sum to 419, but the printed percentages sum to 94.9%. Counts are auditable; treat the printed percentages as paper-printed. |
| As of | 2024-07 |
| Source | Meta AI, The Llama 3 Herd of Models, arXiv v1, July 31, 2024, §3.3.4 / Table 5. · §3.3.4 and Table 5: 466 total interruptions, 47 planned, 419 unexpected; Faulty GPU 148, GPU HBM3 Memory 72, Software Bug 54, Network Switch/Cable 35 |
| Review | checking…review by 2026-12-22 · standard cadence |
| Recorded changes | last 2026-09-16 · 5 revisions tracked |
| Claim id | share-of-training-interruptions-from-faulty |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.