KV-cache memory reduction from attention architecture and quantization; DeepSeek-V2 reported a 93.3% reduction (~15×) from MLA versus its DeepSeek-67B baseline
93.3% (~15x)observed
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | These figures measure KV-cache memory, not end-to-end inference cost. The supported engineering calculation is 135 GB → ~10 GB (~14×). Compression ratios from GQA, MLA, quantization, pruning, and residual coding are not interchangeable and should not be presented as one 4–40× cost-reduction range. |
| Caveat | Deployed DeepSeek-V2 vs DeepSeek 67B. The report attributes this to MLA PLUS inference optimizations including KV-cache quantization to ~6 bits — it is not an MLA-only result. 93.3% leaves 6.7% of baseline, i.e. ~15x. |
| As of | 2024-05 |
| Source | DeepSeek-AI, DeepSeek-V2 Technical Report · Abstract and Table 1 discussion of MLA KV-cache reduction |
| Review | checking…review by 2026-10-01 · fast cadence |
| Recorded changes | last 2026-07-27 · 3 revisions tracked |
| Claim id | long-context-cost-reduction-from-kv-compression |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.