NCCL · NVIDIA Collective Communications Library
NVIDIA's library implementing optimized multi-GPU collective operations like all-reduce over NVLink and the fabric.
Current numbers
ROCm 7.xAMD stack with RCCL NCCL-API parity; MI350X/MI355X support (7.0 Sep 2025)
14.4 TFLOPSin-network compute per SHARPv4 (Quantum-X800), 9x prior gen; FP8, NCCL 2.27 integrated
~16 → ≤6 SMsGPU SMs consumed by a reduction after composing NVLink-SHARP + IB-SHARP in NCCL 2.27