Size a host pool from CPU service demand — stated inputs
100 requests/s; 0.20 CPU-seconds/request; CPU occupancy ceiling 70%; thin host 16 cores; pool 32 cores; 30 active requests × 2.0 GB state; 40 GB runtime reserve; 128 GB host memory; sensitivity 120 requests/s.modeled
| Value kind | modeled — Interpret this value according to its displayed kind, scope, source, and as-of date. |
|---|---|
| Scope | CPU work excludes waiting. Unsupported host workloads, allocations and ranges are explained in the opening callout; Chapter 14.1 owns accounting. Neither CUDA profiling nor occupancy proves these inputs or p99. |
| As of | 2026-09 |
| Source | NVIDIA CUDA best-practices profiling guidance — method reference; guide-authored hypothetical scenario September 8, 2026. No supplier quote or test measurement. |
| Derivation | Exact teaching assumptions chosen to expose the stated constraint and its reversal; not estimates of market prices or actual equipment. |
| Review | checking…review by 2027-03-08 · standard cadence |
| Recorded changes | last 2026-09-16 |
| Claim id | guide2-t29-7-8-inputs |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.