NCCL · NVIDIA Collective Communications Library
NVIDIA's library implementing optimized multi-GPU collective operations like all-reduce over NVLink and the fabric.
Current numbers
ROCm 10.0.0 (release snapshot)AMD stack with RCCL NCCL-API parity; MI350X/MI355X support (7.0 Sep 2025)
~16 → ≤6 SMsGPU SMs consumed by a reduction after composing NVLink-SHARP + IB-SHARP in NCCL 2.27
≈370 of 400 GB/s (~92%) in Together AI's published cluster-testing rule of thumb; the project baseline and declared tolerance set the gateNCCL all_reduce busbw vs node aggregate line rate — Together AI's published rule of thumb; your golden-node baseline sets the gate