reported 2:1–3:1 oversubscription examples in some inference designs; a named 2:1 versus 1:1 model estimates an approximately 31% back-end cost delta
2:1–3:1 examples; ~31% modeled 2:1 cost deltaderivedcontested
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | Examples, not workload defaults. Distributed MoE inference, KV movement, or prefill/decode disaggregation can require high bisection; derive each tier from measured traffic, placement, failure headroom, and the tail-latency SLO. |
| As of | 2025 |
| Source | SemiAnalysis AI Neocloud Playbook |
| Review | checking…review by 2026-09-20 · fast cadence |
| Recorded changes | last 2026-08-24 · 2 revisions tracked |
| Claim id | inference-fabric-oversubscription-vs-1-1-non |
Where the guide uses it
- 1.3 Inference Data Centers: Bursty, Distributed, Always-On
- 10.11 Inference Serving Engineering: SLOs, Batching, Disaggregation & Goodput-Optimal Scheduling
← Full numbers register — every date-stamped figure in the guide, with revision history.