NIXL · NVIDIA Inference Xfer Library
The de facto library for moving KV cache across memory, network, and storage tiers, which is what makes prefill/decode disaggregation practical by turning KV transfer into a scheduled object.
Current numbers
~7x / ~10xDynamo + wide-EP MoE throughput on GB200 NVL72 vs B200 (Dynamo 1.0 GA at GTC 2026); NIXL+GPUDirect Storage prefill speedup for long context