All-reduce
A collective operation that sums gradients across all GPUs and shares the result, the dominant traffic in training.
Current numbers
~2xeffective all-reduce bandwidth from in-network reduction (SHARP) vs non-SHARP config
870-928 GB/sin-domain all-reduce busbw on GB200 NVL72 (~saturates the 900 GB/s/GPU NVLink5 unidirectional rate; NVLS lets ring-convention busbw read slightly above it); scale-out gate set as % of this