more users served per GPU with prefix caching + KV offload combined
~10xestimate
| Value kind | estimate — Interpret this value according to its displayed kind, scope, source, and as-of date. |
|---|---|
| Scope | A secondary summary of a vendor-reported result on one complete tested configuration; the model, serving engine, prefix-hit rate and latency target are unstated. It is not a measured multiplier for another fleet — derive the served-user gain from the prefix-hit rate, the complete transfer time and the remaining time-to-first-token budget, as this chapter does. |
| As of | 2026 |
| Source | Spheron KV-cache optimization guide |
| Review | checking…review by 2026-09-29 · standard cadence |
| Recorded changes | last 2026-06-29 |
| Claim id | more-users-served-per-gpu-with-prefix-caching |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.