GPU- and HBM-attributed interruptions in one Llama 3 405B run: 148 GPU and 72 HBM events among 419 unplanned interruptions over 54 days on 16,384 H100s
148 GPU + 72 HBM interruptionsobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Job-interruption event counts for one named run. The source does not establish unique failed devices or equipment population-time exposure, so these counts must not be annualized into component AFR, fleet lambda, cumulative equipment risk or spares demand. |
| As of | 2024 |
| Source | Meta (The Llama 3 Herd of Models, arXiv 2407.21783) · Table 5, reliability analysis: 419 unplanned interruptions over 54 days; 148 faulty-GPU and 72 HBM3-attributed events |
| Review | checking…review by 2026-09-26 · fast cadence |
| Recorded changes | last 2026-08-24 · 2 revisions tracked |
| Claim id | combined-h100-gpu-hbm-annualized-failure-rate-1 |
Where the guide uses it
- 12.5 Quantitative Reliability & Availability Modeling (RBD / FTA / Monte-Carlo)
- 14.3 Component Failure Modes, Failure Rates & Fleet Reliability Data
← Full numbers register — every date-stamped figure in the guide, with revision history.