SHARP · Scalable Hierarchical Aggregation and Reduction Protocol
In-network computing that lets switches perform parts of a collective (such as summing gradients) so data crosses the fabric fewer times.
Current numbers
NCCL 2.28+NVLink-SHARP Multimem multicast + device-side communication API within an NVL72 domain
~2xeffective all-reduce bandwidth from in-network reduction (SHARP) vs non-SHARP config
~16 → ≤6 SMsGPU SMs consumed by a reduction after composing NVLink-SHARP + IB-SHARP in NCCL 2.27