unplanned interruptions on 16,384 H100s (~1 / 3 hr); 78% hardware-caused
419 / 54 daysobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability. |
| Caveat | The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator. |
| As of | 2024 |
| Source | Meta Llama 3 paper (Table 5) |
| Review | checking…review by 2026-09-03 · standard cadence |
| Recorded changes | last 2026-06-29 |
| Claim id | unplanned-interruptions-on-16-384-h100s-1-3-hr |
Where the guide uses it
- 1.2 Training Data Centers: Synchronous, Dense, Checkpointable
- 10.6 Observability, Telemetry & GPU Health
- 10.7 Fleet Reliability, Fault Tolerance & Autonomous Recovery
- 12.1 Resilience Standards, Redundancy Topologies & Fault-Domain Engineering
- 12.2 The AI-Cluster Reliability Rethink: Goodput vs Facility Availability
- 13.1 Commissioning Fundamentals, Levels & Program Governance
- 14.1 Operational KPIs, Goodput & the Reliability Economics of AI Factories
- 14.4 Reliability Engineering for Training (Operational)
← Full numbers register — every date-stamped figure in the guide, with revision history.