Together AI published approximately 370 of 400 GB/s as a cluster-testing rule of thumb; not a portable NCCL acceptance threshold
≈370 of 400 GB/s (~92%) in Together AI's published cluster-testing rule of thumb; the project baseline and declared tolerance set the gateestimate
| Value kind | estimate — Interpret this value according to its displayed kind, scope, source, and as-of date. |
|---|---|
| Scope | Historical rule of thumb only. The project gate requires a declared collective, message size, ranks, library/transport/algorithm, metric normalization, topology and measured baseline tolerance. |
| As of | 2024-08-13 |
| Source | Together AI, “A practitioner’s guide to testing and running large GPU clusters for training generative AI models,” 13 August 2024; practitioner 370/400 example. · 'looking for the all_reduce_perf test to show bandwidth around 92% of the theoretical maximum of the fabric: so around 370 GB/s on a 400GB/s fabric' |
| Derivation | 370 / 400 = 0.925. The source does not specify a transferable NIC-count, rank-count or normalized-busbw conversion; do not infer an eight-rail configuration. |
| Review | checking…review by 2026-11-23 · standard cadence |
| Recorded changes | last 2026-09-16 · 5 revisions tracked |
| Claim id | nccl-all-reduce-busbw-vs-theoretical-acceptance |
Where the guide uses it
- 10.4 Node Software Stack: Drivers, CUDA/ROCm, NCCL & Firmware
- 10.5 Provisioning, Bring-Up & Infrastructure as Code
← Full numbers register — every date-stamped figure in the guide, with revision history.