SHARP · Scalable Hierarchical Aggregation and Reduction Protocol
In-network computing that lets switches perform parts of a collective (such as summing gradients) so data crosses the fabric fewer times.
Current numbers
NCCL 2.28+NVLink-SHARP Multimem multicast + device-side communication API within an NVL72 domain
~16 → ≤6 SMsGPU SMs consumed by a reduction after composing NVLink-SHARP + IB-SHARP in NCCL 2.27