share of GPU VRAM the KV-cache consumes at 1M-token context
70–90%estimate
| Value kind | estimate — Interpret this value according to its displayed kind, scope, source, and as-of date. |
|---|---|
| Scope | A secondary-source range for 1M-token context, not a measurement of a named deployment. KV share follows from layers, KV-head count, head dimension, precision and per-rank sharding; derive it for the served model and its parallelism layout rather than importing this band. |
| As of | 2026 |
| Source | Spheron / DigitalApplied KV optimization guides |
| Review | checking…review by 2026-08-16 · standard cadence |
| Recorded changes | last 2026-06-29 |
| Claim id | share-of-gpu-vram-the-kv-cache-consumes-at-1m |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.