Google multi-tier checkpointing: 6.59% goodput uplift on a 35K-chip TPU v5p workload, checkpoint-save latency under five minutes, and restore under one minute across thousands of nodes
+6.59% goodput; save <5 min; restore <1 minobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Google reports the figure as a 6.59% increase in ML Goodput on one 35K-chip TPU v5p workload; that wording does not establish 6.59 percentage points absolute. The save and restore latencies are checkpoint-path measurements, not an end-to-end detection-to-resumed-training MTTR for an arbitrary GPU fleet. |
| As of | 2025 |
| Source | Google Cloud (multi-tier checkpointing) |
| Review | checking…review by 2027-05-01 · standard cadence |
| Recorded changes | last 2026-08-31 · 2 revisions tracked |
| Claim id | mttr-reduction-from-persistent-storage-reload |
Where the guide uses it
- 10.7 Fleet Reliability, Fault Tolerance & Autonomous Recovery
- 12.2 The AI-Cluster Reliability Rethink: Goodput vs Facility Availability
- 14.4 Reliability Engineering for Training (Operational)
← Full numbers register — every date-stamped figure in the guide, with revision history.