observed cadence for one 16,384-H100 Llama 3 405B run: 419 unplanned interruptions over 54 days
every ~3 hrobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Whole-job events for the named Meta run, including hardware and software causes. Do not transform by GPU count or interpret as per-GPU MTBF. |
| As of | 2024 |
| Source | Meta Llama 3 405B disclosure |
| Review | checking…review by 2026-11-23 · standard cadence |
| Recorded changes | last 2026-06-29 |
| Claim id | failure-cadence-of-a-16k-gpu-cluster-llama-3 |
Where the guide uses it
- 10.1 Orchestration Architecture & the Scheduling Plane
- 10.4 Node Software Stack: Drivers, CUDA/ROCm, NCCL & Firmware
← Full numbers register — every date-stamped figure in the guide, with revision history.