observed interruption attribution in one Llama 3 405B run: 148 faulty-GPU and 72 HBM3 events among 419 unplanned interruptions
148 GPU + 72 HBM interruptionsobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Table 5's printed percentages are internally inconsistent; retain the event counts and the named 16,384-H100, 54-day job boundary. Counts are not unique failed FRUs or equipment AFR. |
| Caveat | Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown. |
| As of | 2024-07 |
| Source | Meta AI, The Llama 3 Herd of Models, §3.3.4/Table 5 · §3.3.4, Table 5: 466 total interruptions, 47 planned, 419 unexpected; rows for faulty GPU, HBM3, software, network, and the remaining categories |
| Review | checking…review by 2026-12-22 · standard cadence |
| Recorded changes | last 2026-08-24 · 3 revisions tracked |
| Claim id | share-of-interruptions-gpu-hbm-related-gpu-30-1 |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.