Quantization
Using lower numerical precision (FP8, FP4, INT8) to cut memory and boost throughput, trading some accuracy for speed.
Current numbers
93.3% (~15x)KV-cache memory reduction from attention architecture and quantization; DeepSeek-V2 reported a 93.3% reduction (~15×) from MLA versus its DeepSeek-67B baseline