NIXL · NVIDIA Inference Xfer Library
The de facto library for moving KV cache across memory, network, and storage tiers, which is what makes prefill/decode disaggregation practical by turning KV transfer into a scheduled object.
The de facto library for moving KV cache across memory, network, and storage tiers, which is what makes prefill/decode disaggregation practical by turning KV transfer into a scheduled object.