power-oversubscription headroom: inference (uncorrelated peaks) vs synchronous training
~21% vs ~3%observed
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Measured on the production LLM fleets and power-management platform characterized by POLCA (Patel et al., ASPLOS 2024). It is not a deployable allowance for a 2026 reasoning, MoE, or disaggregated fleet: derive that from the proposed fleet's measured coincident demand, its tested protection and capping response, and the serving degradation you will allow. |
| As of | 2024 |
| Source | Patel et al., POLCA (Microsoft Research), ASPLOS 2024 · Abstract: training clusters offer about 3% headroom, versus about 21% for inference clusters |
| Review | checking…review by 2026-11-23 · standard cadence |
| Recorded changes | last 2026-07-03 · 2 revisions tracked |
| Claim id | power-oversubscription-headroom-inference-2 |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.