Collective
A coordinated communication pattern across many GPUs (all-reduce, all-gather, reduce-scatter) central to distributed AI.
Current numbers
1.8 → 3.6 TB/sNVLink Gen5 Blackwell / Gen6 Rubin per-GPU bidirectional ratings; NVL72 sums 72×1.8=129.6 and 72×3.6=259.2 TB/s (~130/~260), not delivered collective throughput
870-928 GB/sin-domain all-reduce busbw on GB200 NVL72 (~saturates the 900 GB/s/GPU NVLink5 unidirectional rate; NVLS lets ring-convention busbw read slightly above it); scale-out gate derived independently from NICs, rails, topology, protocol efficiency, message size, and collective model