Burn-in
Running new hardware hard for a period to surface early ('infant mortality') failures before production use.
Current numbers
~1 failure / 512 GPUs / weekbest-in-class fleet failure rate after burn-in; new clusters fail far more for the first 3–4 weeks — the bring-up tail
~7 daysbest-in-class H100 MTBF per 512 GPUs in a mature cluster; new clusters far worse (3–4 wk burn-in)
72–168 hrtypical burn-in soak before a new cluster is admitted to production