Chapter 9.7
In this chapter · 7 sections
Inference & KV-Cache Storage: The New Memory Hierarchy
Inference made the KV-cache a first-class storage problem: its bytes now spill past HBM, and where you let them land — DRAM, CXL, NVMe, or Ethernet-flash — sets your tokens-per-second and cost-per-token.
What you'll decide here
- Whether your inference fleet treats the KV-cache as ephemeral HBM state (recompute on miss) or as a managed, multi-tier asset that is offloaded, persisted, and reused across requests and sessions.
- Where the offload tier physically lives — node-local DRAM, CXL-expanded/pooled DRAM, local NVMe, or a network-attached Ethernet-flash KV tier — and the latency/bandwidth/cost envelope each choice imposes on time-to-first-token.
- The rule for when CXL-expanded DRAM beats NVMe KV-offload (and when it does not): a capacity-and-reuse-rate decision, not a vendor preference.
- Whether to adopt a KV-transfer and reuse stack (Dynamo/KVBM + NIXL, LMCache, vLLM PagedAttention, Mooncake-class pools) or recompute, and how that choice couples to prefill/decode disaggregation and the back-end fabric.
- How model/weight serving and cold-start are tiered alongside the KV-cache — because the same memory hierarchy that holds context also gates how fast a model loads onto a freed GPU.
For a decade the storage conversation in AI was about training — feed the GPUs fast enough (ingestion), and don't lose the run (checkpointing). Inference was assumed to be storage-light: load the weights once, serve from HBM, done. That assumption is dead. The workload Deloitte’s November 2025 outlook forecasts at roughly two-thirds of 2026 AI compute (Chapter 1.3) created a new storage tier with no precedent in the training stack: the KV-cache — the per-token attention state every autoregressive request must keep resident for as long as it is generating. Reasoning models that emit tens of thousands of decode tokens, agents that carry long histories, and RAG pipelines with multi-thousand-token system prompts have turned that state from a rounding error into the dominant consumer of the most expensive memory in the building.
This chapter treats the KV-cache as what it has become: a memory-hierarchy problem with a storage answer. We walk the hierarchy from HBM down through local DRAM, CXL-expanded DRAM, NVMe, Ethernet-attached flash, and object storage; we make the present-tense case for CXL as a deployable tier and for when it beats NVMe; and we cover the KV offload-and-reuse machinery — Dynamo/KVBM, NIXL, LMCache, prefix caching — plus the model-weight-serving and cold-start problem that rides the same hierarchy. Every tier you add buys capacity and reuse at the price of latency on the miss path, and the wrong placement shows up directly as a blown time-to-first-token SLO.
Why inference created a new tier
Start with the tensor layout. Logical KV bytes for a dense attention cache are 2 × layers × KV heads × head dimension × bytes per element × retained tokens × sequences. The leading two counts keys and values. Grouped-query attention uses the KV-head count, not the query-head count. This is aggregate state across participating ranks; derive each rank's actual share from tensor, pipeline and context parallelism, including any replicated KV heads. Then add allocator, alignment and layout overhead and compare it with HBM remaining after the local weights and workspaces. The worked prefix case uses Meta's Table 3 dimensions and a declared precision; model parameters alone cannot determine KV bytes.
HBM is scarce and expensive, but the aggregate sequence figure alone does not prove that a deployment must spill. First derive the actual per-rank layout, replication and overhead for the named tensor-, pipeline-, context- and data-parallel configuration, then compare it with usable per-rank HBM. If the allocation does not fit the service's batch, context and latency envelope, the choices include architectural compression (GQA or MLA), FP8/FP4 KV quantization, lower concurrency, recomputation, or offload to a slower tier. Offload is then a tiered-storage decision on the request's critical path, with transfer bandwidth, reuse and tail latency measured for the workload.
The hierarchy: HBM → DRAM → CXL → NVMe → Ethernet-flash → object
The new memory hierarchy is a ladder of declining cost and rising latency. Each rung holds colder, larger, cheaper KV state than the one above it. The engineering job is to keep the hottest, most-reused blocks high and let everything else fall — and to size the bandwidth between rungs so that a promotion on a hit does not blow the budget the hit was supposed to save.
NVIDIA's BlueField-4 Context Memory Storage Platform (CMX, formerly ICMSP) formalized this ladder into named tiers — G1 (GPU HBM), G2 (host CPU DRAM), G3 (node-local NVMe), a new G3.5 (an Ethernet-attached flash tier optimized specifically for KV-cache), and G4 (external durable NVMe) — and it is the clearest sign that the industry now treats KV-cache placement as a storage-architecture decision, not an inference-engine implementation detail (NVIDIA / Blocks & Files, 2026). The table below is the decision surface.
| Tier | Medium | Access latency | Bandwidth (per device) | Capacity / cost | Role in KV serving |
|---|---|---|---|---|---|
| G1 | GPU HBM (HBM3E / HBM4) | ~100 ns | ~1.2+ TB/s per HBM3E stack, ~2 TB/s per HBM4 stack; ~5–8 TB/s aggregate across a current GPU's stacks | Tiny / extreme $/GB | Live working set: active decode KV + weights |
| G2 | Host CPU DRAM (DDR5) | ~80–140 ns local | ~40–50 GB/s per channel | Modest / high $/GB | First spill tier; warm reusable blocks |
| CXL | CXL-expanded / pooled DRAM | ~250–600 ns (load/store) | ~tens of GB/s per link | Large, memory-semantic | Byte-addressable capacity tier; pooled reuse |
| G3 | Node-local NVMe (TLC/QLC) | ~10–100 µs | ~6–14 GB/s (PCIe 5.0; ~28 GB/s at 6.0) | Large / low $/GB | Cost-optimized prefix/session cache |
| G3.5 | Ethernet-attached flash (CMX) | ~100 µs class | 320 GB/s read class (WEKA/STX) | Very large / shared | Networked KV tier across the fleet |
| G4 / cold | External NVMe / object store | ms+ | Throughput-bound | Effectively unbounded / cheapest | Durable context, persisted sessions, model store |
The hierarchy is a set of cliffs, not a smooth gradient. The two that matter most: the jump from DRAM/CXL (nanoseconds, byte-addressable, load/store) to NVMe (microseconds, block-addressable, DMA) is a latency step from memory access to queued storage I/O and bulk transfer — that is the line between "memory" and "storage," and it is where the programming model changes. The jump from node-local to network-attached (G3 to G3.5/G4) trades single-node capacity for fleet-wide reuse, at the price of a fabric round-trip — which can buy either shared reuse or capacity beyond one node. Measure both the saved compute and the transfer contention before choosing it. Place a hot block one cliff too low and the latency of a hit approaches that of a miss; place a cold block one cliff too high and you have evicted a block someone else needed to make room for one nobody will reuse.
CXL DRAM as a present-tense tier
By 2026 CXL is a deployable tier with two distinct uses that get conflated and should not be. Memory expansion adds byte-addressable DRAM to a single host beyond its DIMM-channel limit, over the PCIe physical layer with cache-coherent load/store semantics — the OS and the inference engine see more memory, full stop. Memory pooling disaggregates DRAM into a capacity resource that many hosts draw allocations from — flexible assignment, not shared bytes. CXL 2.0 pooling gives each region to one host at a time; concurrent multi-host access to the same physical KV blocks is memory sharing, a CXL 3.x capability that needs matching switch, device, and software support. State the generation and topology before you assume cross-server prefix reuse, because the two buy very different things. Commercial CXL pools reached 100 TiB scale in 2025 with larger 2026 deployments, and Samsung's CMM-D-in-a-CXL-switch work is explicitly pitched at KV-cache offload (Samsung CMM-D white paper, June 2026).
The decisive property is that CXL keeps the cache memory-semantic. NVMe offload forces a block-mode DMA round-trip and a copy; CXL-expanded DRAM is a host load away. That last phrase is a CPU-side property, and the consumer here is a GPU: budget the whole path — NUMA placement, the host-to-device hop, and the runtime support that has to exist — rather than reading CXL load latency as the accelerator's fetch time. The published evidence for how much that is worth is thin. The speedup and GPU-memory-saving multipliers circulating for CXL KV pools do not trace back to a paper with a stated configuration, baseline path, workload, and latency target, so treat them as unestablished and measure the fetch path on your own stack before you buy to them. FMS 2026 added a measured near-memory rung: Marvell's Structera A with SK hynix CMM-Ax offloads KV overflow to high-capacity CXL DRAM and processes it in place — vendor-reported up to 5.5× throughput vs a single-GPU baseline on Llama-3-8B at 1M context (one model, one vendor stack). Marvell also proposed a pod-level optical shared-memory tier ("Photonic Fabric": up to 32 TB of warm KV at 50 m reach) — a watch item, not a deployable 2026 SKU. For scale context on the Ethernet-flash rung: a Vera Rubin SuperPOD-scale CMX complex pencils out at ~9.6 PB of flash (16 enclosures × 4 BlueField-4 × 150 TB), which is why the KV tier is now a NAND-procurement line as well as a latency class. The cost is real estate and complexity: CXL load latency (~250–600 ns) is several times host-local DRAM, the controllers and switches are a new BOM line, and pooling needs a memory-fabric topology that most halls were not built for (Chapter 8.5 for the fabric framing).
One token requires 2 × 80 × 8 × 128 × 2 = 327,680 bytes aggregate. A full sequence therefore occupies 327,680 × 131,072 = 40 GiB aggregate; ideal head sharding gives 5 GiB/rank. With 20% overhead, each full sequence needs 6.0 GiB/rank: floor(60/6.0) = 10 sequences. Eleven require 66 GiB and fail capacity before any throughput claim. Reserve growth beyond the stated retained-token ceiling separately.
The reusable prefix occupies 327,680 × 8,192 = 2.5 GiB aggregate, or 320 MiB/rank = 0.33554432 GB/rank. One fetch takes 0.33554432 GB / 20 GB/s + 0.0050 s ≈ 22 ms, within 50 ms and faster than the assumed 150 ms recompute. Select fetch for this authorized prefix, preserving the budget for its uncached suffix. The result earns a transfer choice, not a user-count multiplier.
Flip: two simultaneous promotions need 2 × 0.33554432/20 + 0.0050 ≈ 39 ms and pass; three need about 55 ms and fail. The crossover for three is 3 × 0.33554432/(0.050 − 0.0050) ≈ 22 GB/s, so the assumed 20 GB/s path admits at most two promotions under this sharing model. Schedule earlier prefetch, cap simultaneous misses or add delivery capacity. Recompute at 150 ms cannot rescue this deadline. Qualify queue behavior and all participating ranks at that admission limit, then hand the allowed concurrency to Chapter 1.3 and transfer demand to Chapter 8.5.
Scope & caveats
Illustrative inputs; qualify project rates and limits before selection.
KV-cache offload, reuse, and the transfer fabric
A hierarchy is inert without a manager that moves blocks across it and decides what to keep. In 2026 that machinery converged on a recognizable stack. PagedAttention (vLLM) made the KV-cache a paged, non-contiguous structure — the precondition for everything else, because you cannot offload or share a cache you cannot address in blocks. On top of it, prefix caching reuses the KV of a shared prompt prefix across requests: for RAG and agent workloads with long, repeated system prompts, the prefix's blocks stay resident instead of being recomputed per request, and the saved prefill work can increase service capacity only while prefix transfer, suffix computation and decode remain within the admitted workload's SLO.
The transport layer is where the storage and networking disciplines fuse. NVIDIA Dynamo introduced the KV Block Manager (KVBM), which decouples KV memory management from the inference engine and orchestrates movement across G1–G4; it uses NIXL (the NVIDIA Inference Transfer Library) as the unified transport across NVLink, RDMA NICs, and GPUDirect Storage, and integrates LMCache for reuse and eviction. Dynamo, NIXL and GDS occupy different parts of the serving stack: the manager decides KV residency, the transfer layer moves buffers, and the execution engine defines compatible KV layout and consumes it. Pin all three releases and measure the complete request path on the selected model. A GB200 NVL72 versus B200 comparison that also changes expert parallelism cannot isolate a storage-reuse gain; neither can an unqualified prefill multiplier. NIXL itself stopped being NVIDIA-only plumbing in mid-2026: NIXL 1.3 (Jul 2026) ships DDN's Infinia plugin inside the standard wheel and the prebuilt Dynamo container — the first storage vendor in the official package — and AMD's RIXL fork was merged upstream on 2026-06-04 and deprecated, after validating end-to-end VRAM transfers on MI300X/MI355X (SemiAnalysis). The transfer library is becoming the inference data-movement ABI that storage vendors and the second GPU vendor target, which softens the certified-appliance lock-in calculus in Chapter 9.2. Open alternatives — LMCache standalone, Mooncake-class disaggregated KV pools, llm-d — chase the same pattern. The architectural consequence is that KV reuse pulls you toward prefill/decode disaggregation: separate the compute-bound prefill pool from the bandwidth-bound decode pool, generate KV once in prefill, and transfer it to decode over the fabric, which makes the KV-transfer path a primary fabric-sizing input, not an afterthought (Chapter 8.5; Chapter 9.3 for the GPUDirect Storage data path).
Deep dive: recompute vs offload — the decision the KV manager makes thousands of times a second
On every cache miss the serving stack faces the same micro-decision: recompute the missing KV (run prefill again, spend GPU FLOPs) or fetch it from a lower tier (spend transfer latency and bandwidth). The crossover is a function of three numbers: the prefill cost of the prompt (FLOPs, which grow with prefix length), the fetch latency of the tier holding the block, and the bandwidth available to move it. For a short prompt on a busy GPU, recompute is cheaper than a fabric round-trip — the block is small and the GPU has spare prefill capacity at low concurrency. For a long shared prefix (an illustrative 4,000-token RAG system prompt reused across thousands of requests), repeated prefill wastes compute when fetching compatible cached KV — even from G3.5 over the network — meets the deadline faster, provided the compatible-prefix transfer and uncached suffix actually fit the remaining deadline.
The trap is treating this as a static config. The right answer flips with concurrency: as the decode pool fills and prefill capacity becomes the bottleneck, the recompute option gets more expensive (it competes with paying users for FLOPs) and offload gets relatively cheaper, so the crossover point moves toward fetch. A KV manager that decides once at deploy time leaves goodput on the table at both ends; managers such as KVBM and LMCache need their admission and eviction behavior qualified per supported block and tier against live queue depth; verify whether the selected release actually makes the recompute-versus-fetch decision you need. This is the inference analogue of the checkpoint-interval optimization in Chapter 9.4 — a cost-balanced decision, made continuously instead of once.
Scope & caveats
Aggregate logical KV at stated precision; excludes weights, workspaces, allocator overhead and replication. Derive actual per-rank placement before testing fit.
Scope & caveats
Aggregate logical KV at stated precision; excludes weights, workspaces, allocator overhead and replication. Derive actual per-rank placement before testing fit.
Scope & caveats
A secondary-source range for 1M-token context, not a measurement of a named deployment. KV share follows from layers, KV-head count, head dimension, precision and per-rank sharding; derive it for the served model and its parallelism layout rather than importing this band.
Scope & caveats
These figures measure KV-cache memory, not end-to-end inference cost. The supported engineering calculation is 135 GB → ~10 GB (~14×). Compression ratios from GQA, MLA, quantization, pruning, and residual coding are not interchangeable and should not be presented as one 4–40× cost-reduction range.
Deployed DeepSeek-V2 vs DeepSeek 67B. The report attributes this to MLA PLUS inference optimizations including KV-cache quantization to ~6 bits — it is not an MLA-only result. 93.3% leaves 6.7% of baseline, i.e. ~15x.
Scope & caveats
A secondary summary of a vendor-reported result on one complete tested configuration; the model, serving engine, prefix-hit rate and latency target are unstated. It is not a measured multiplier for another fleet — derive the served-user gain from the prefix-hit rate, the complete transfer time and the remaining time-to-first-token budget, as this chapter does.
Model and weight serving: the cold-start tax on the same hierarchy
The KV-cache is not the only thing riding this hierarchy. Multi-model and autoscaled inference fleets must also load weights onto GPUs that were just freed — and the same tiers govern how fast that happens. A 70B model in FP8 is ~70 GB; loading it from a slow object store across the network is a multi-minute stall during which a freshly-scaled GPU earns nothing. This is the cold-start tax, and it is the mirror image of the KV problem: where KV is about keeping per-request state warm, weight serving is about keeping per-model state reachable fast enough to scale.
The placement logic is the same ladder applied to weights. The hot path keeps the working set of frequently-served models on node-local NVMe (G3) or a fast network-flash tier so a scale-up event is a fast block read, not a cold object pull; the cold path keeps the long tail of rarely-served models and durable artifacts in object storage (Chapter 9.6, which treats object as the inference cold tier and model-distribution backbone). Under-provision the weight-serving tier and the consequence is concrete: your autoscaler's response time is gated by storage, not by GPU availability — you scale slower than your traffic spikes, breach SLOs during the ramp, and over-provision idle GPUs to compensate. The KV hierarchy and the weight hierarchy are the same physical media, contending for the same bandwidth, and they must be sized together.
Multi-tier KV management and the placement policy
With the tiers, the transport, and the workloads in place, the remaining decision is policy: what gets promoted, what gets evicted, and where each block lands. This is where most of the realized goodput is decided, because the hardware ladder only sets the ceiling; the policy determines how close you get. Three policy axes matter. Eviction: LRU is the default, but reuse-aware policies that weight by prefix-sharing frequency keep high-fanout system prompts resident far longer than a naive recency scheme would, which is the whole point of prefix caching. Placement: the manager must route each block by modeled fetch time, hit probability, and prefetchability against the remaining TTFT and recompute budgets — the CXL-vs-NVMe rule above, applied per block rather than per fleet. Coherence and sharing: a pooled or networked tier (CXL pool, G3.5 Ethernet-flash) can extend capacity or let authorized GPUs reuse compatible context; CXL pooling by itself allocates capacity and does not establish simultaneous sharing. Genuine sharing is enormously efficient for shared prefixes, but it introduces a consistency and lifetime-management problem that node-local caches never had. Get the policy wrong and the symptoms are unambiguous: thrashing between tiers (blocks promoted and evicted before reuse), cache pollution (cold blocks crowding out warm ones), or a TTFT tail that tracks the slowest tier instead of the fastest hit. The fleet-level sizing of these tiers — how much DRAM, CXL, NVMe, and network-flash per GPU, and the bandwidth between them — is the storage:compute co-design problem taken up in Chapter 9.8.
Key reuse by the effective model/adapter revision, token IDs, position convention, KV dtype/layout and authorized sharing group. Prefix identity must include the preceding context; equal trailing tokens alone are insufficient. With the selected runtime, issue an allowed repeated prefix and verify reuse plus equivalent output. Change model revision, adapter, token prefix or tenant sharing key and require a miss. Revoke a session and confirm its cached state stops being served under the declared retention policy. vLLM prefix-cache identity and salting provide one implementation; test the configured behavior rather than assuming a cache hit is permission.
Deep dive: why disaggregated inference makes the KV transfer a fabric-design problem
Prefill and decode have opposite hardware appetites. Prefill is compute-bound: it processes the whole prompt in parallel and saturates GPU FLOPs. Decode is memory-bandwidth-bound: it generates one token at a time and is starved by HBM bandwidth, not FLOPs. Running both on the same GPU means each phase under-utilizes the resource the other needs. Disaggregated inference splits them into separate pools — a prefill pool that builds the KV-cache and a decode pool that consumes it — so each can be sized and scaled independently. NVIDIA's GB200 NVL72 + Dynamo work is built around exactly this split, transferring KV over NVLink between the pools.
The consequence lands in the fabric. Disaggregation means the KV-cache produced in prefill must move to wherever decode runs, on every request, on the critical path. The transfer is large (gigabytes for long context) and latency-sensitive (it sits in front of the first decode token), so the link between prefill and decode pools becomes a primary fabric-sizing constraint — size the GPU-to-GPU path, NVLink in the cited in-rack design, to B_required ≥ concurrent_KV_bytes / transfer_window, then add headroom for p99 efficiency loss, protocol overhead, and oversubscription; size storage-to-GPU prestaging over RDMA against its separate transfer window, or disaggregation injects TTFT latency. This is why NIXL spans NVLink, RDMA, and GPUDirect Storage with one API: the KV transfer may ride any of them depending on where the tiers sit, and the fabric must be co-designed to carry it. Topology and oversubscription for this traffic are engineered in Chapter 8.5; the CPU-bypass data path that makes storage-to-GPU KV streaming viable is Chapter 9.3.
Vectors occupy 100,000,000 × 768 × 4 = 307.2 GB; one index needs 307.2 + 64 + 32 = 403.2 GB. Two replicas with an old/new rebuild need 403.2 × 2 × 2 = 1,612.8 GB; at 80% occupancy, provision 1,612.8/0.8 ≈ 2.0 TB. The 2.1 TB reservation passes this assumed footprint. Flip: eight-byte vectors require (614.4 + 64 + 32) × 4/0.8 ≈ 3.6 TB, so the rebuild fails capacity. The numerical crossover is about 4.2 bytes/element at the same allowances; supported precision choices and index compression must also preserve retrieval quality.
Candidate selection remains HOLD until a labeled query set passes both assumed gates. Use exact search over the same authorized corpus as the filtered-recall reference; report recall by tenant/filter class and query latency at admitted concurrency. Pin embedding model, normalization, distance metric, chunking and corpus version. Delete a source chunk and revoke its permission, then query every serving replica and result cache: forbidden content must cease reaching the prompt by the declared propagation deadline. Return index generation and deletion watermarks with the evidence. This chapter owns the index storage contract; Chapter 10.10 owns data eligibility and Chapter 1.3 owns the request SLO.
Scope & caveats
Illustrative inputs; qualify project rates and limits before selection.
Anti-patterns
The recurring mistakes all come from importing a training-era storage mental model into an inference fleet, or from treating the KV-cache as an implementation detail instead of an architecture:
- Recompute-everything by default. Disabling or under-sizing KV offload because "storage is slow" — then paying full prefill FLOPs on every shared-prefix request. For RAG and agent workloads this is a leading source of wasted inference compute, and it scales with how good your traffic is (more repeated prompts = more waste).
- One tier too low for the transfer budget. Placing a multi-gigabyte warm prefix on NVMe based on media latency alone, then discovering that access latency, bytes divided by effective bandwidth, and queueing blew the TTFT SLO. The fix is to compare complete fetch time with the remaining TTFT budget and recompute time at admitted concurrency before buying more GPUs; add capacity only where that comparison shows it removes the limiting resource.
- Sizing the weight-serving tier as an afterthought. Provisioning HBM and fabric for steady state, then watching the autoscaler stall on cold model pulls from object storage during every traffic spike. Cold-start is a storage-bandwidth problem when weight transfer dominates, with verification, deserialization and warm-up added to the clock; size G3 or stage weights ahead, or the fleet scales slower than its traffic.
- Shared KV across an untrusted boundary. Turning on cross-request prefix sharing or session persistence for the throughput win without scoping it to a trust boundary — converting a performance tier into a multi-tenant data path. → Chapter 11.5.
Choose KV placement from per-rank capacity and the complete transfer deadline, with admission limits that survive simultaneous misses. Keep retrieval indexes, source chunks and model weights in the same physical demand ledger while testing their distinct correctness contracts. A cache hit that arrives late, belongs to another sharing group or uses incompatible model state has no serving value; a larger tier pays only when it increases authorized requests completed within the SLO.
Cite this chapter
Fehn, J. (2026). Inference & KV-Cache Storage: The New Memory Hierarchy (Chapter 9.7). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-9-storage-and-data/9-7-inference-and-kv-cache-storage-the-new-memory-hierarchy (accessed 2026-09-29).
@misc{aidc-9-7,
author = {Fehn, Jacob},
title = {Inference & KV-Cache Storage: The New Memory Hierarchy (Chapter 9.7)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-9-storage-and-data/9-7-inference-and-kv-cache-storage-the-new-memory-hierarchy},
note = {Accessed 2026-09-29}
}