The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

Training RPO is the lost-work age of the state captured by the newest valid checkpoint that survives the declared failure domain

age of newest surviving captured statederived

Value kindderived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method.
ScopeAt failure t, select a valid checkpoint whose durable commit completed and whose storage, keys and metadata survive the declared event. Use its captured step/state time, not its commit completion timestamp, to calculate lost work.
As of2026-08-24
SourcePyTorch documents checkpoint persistence; the distinction between captured progress and commit completion is a guide derivation. · Distributed checkpoint save/async_save completion and state-dict semantics; guide lost-work definition
DerivationLet j be the checkpoint with the newest captured progress among valid surviving checkpoints committed by failure t. Lost work is progress(t) − captured_progress(j); wall-clock state age is t − capture_time(j), not t − commit_time(j).
Reviewchecking…review by 2027-09-05 · standard cadence
Recorded changeslast 2026-09-16 · 5 revisions tracked
Claim idtraining-rpo-floor-set-by-checkpoint-interval

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.