large-LLM job failure rate, top-5% most resource-intensive tasks (Alibaba Unicron); ~37% hardware-attributed, ~73% restart-recoverable
~43.4%observed
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Failure rate of the top-5% most resource-intensive large-LLM jobs in one production fleet (Alibaba, Unicron paper), not a per-job, per-GPU or fleet-wide rate. ~37% of the failures are hardware-attributed and ~73% are recoverable by restart on that fleet; the paper does not establish device-population exposure, so this must not be annualized into a component failure rate or applied to another operator's job mix. |
| As of | 2024 |
| Source | Unicron (He et al., arXiv 2401.00134) |
| Review | checking…review by 2026-08-29 · fast cadence |
| Recorded changes | last 2026-07-03 · 2 revisions tracked |
| Claim id | large-llm-job-failure-rate-top-5-most-resource |
Where the guide uses it
- 10.6 Observability, Telemetry & GPU Health
- 10.7 Fleet Reliability, Fault Tolerance & Autonomous Recovery
- 10.8 MLOps & Training Frameworks
- 12.2 The AI-Cluster Reliability Rethink: Goodput vs Facility Availability
← Full numbers register — every date-stamped figure in the guide, with revision history.