observed interruption interval for one 16,384-H100 Llama 3 405B run: 419 unplanned interruptions over 54 days; why provisioning is a continuous day-2 loop
~3 hrobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Observed whole-job cadence for the named run, not a per-GPU MTBF or portable accelerator-count scaling law. |
| As of | 2024 |
| Source | Meta Llama 3 |
| Review | checking…review by 2026-09-02 · standard cadence |
| Recorded changes | last 2026-08-24 · 2 revisions tracked |
| Claim id | failure-interval-for-a-16-000-gpu-cluster-at-50 |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.