Training RPO is the lost-work age of the state captured by the newest valid checkpoint that survives the declared failure domain
age of newest surviving captured statederived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | At failure t, select a valid checkpoint whose durable commit completed and whose storage, keys and metadata survive the declared event. Use its captured step/state time, not its commit completion timestamp, to calculate lost work. |
| As of | 2026-08-24 |
| Source | PyTorch documents checkpoint persistence; the distinction between captured progress and commit completion is a guide derivation. · Distributed checkpoint save/async_save completion and state-dict semantics; guide lost-work definition |
| Derivation | Let j be the checkpoint with the newest captured progress among valid surviving checkpoints committed by failure t. Lost work is progress(t) − captured_progress(j); wall-clock state age is t − capture_time(j), not t − commit_time(j). |
| Review | checking…review by 2027-09-05 · standard cadence |
| Recorded changes | last 2026-09-16 · 5 revisions tracked |
| Claim id | training-rpo-floor-set-by-checkpoint-interval |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.