A vLLM technique that manages the KV cache in fixed pages like virtual memory, cutting waste and fragmentation.
Related terms: KV cache · vLLM
← All terms