Each box holds one chunk, and the number inside is its value. Ranks pass chunks to their right-hand neighbour, N−1 times. Watch the numbers: in a reduction they add up, in a gather they only change hands — and the box never changes size either way.
The dots say whose numbers have been added in — four dots means the chunk is fully
reduced. They are not a size: 10 + 20 + 30 + 40 = 100 is still one number
occupying one box.
Counted in multiples of the whole buffer. One hop moves one chunk = 1/N of the buffer.
2(N−1)/N. This is the DDP gradient sync.
1/N the size of the input. FSDP and ZeRO use this: a rank owns one shard of
the parameters, so it only needs the summed gradient for that shard.
alg_bw = size / time answers "how fast did my call finish", but it can't be
compared across collectives — the same buffer puts different amounts of traffic on the wire.
Multiplying by the factor gives bus_bw, the traffic the link actually carried,
which is comparable both between collectives and against the hardware's rated bandwidth.
| Collective | Hops per rank | Factor | At N = 32 |
|---|---|---|---|
| all_reduce | 2(N−1) | 2(N−1)/N | 1.9375× |
| reduce_scatter | N−1 | (N−1)/N | 0.9688× |
| all_gather | N−1 | (N−1)/N | 0.9688× |
| all_to_all | N−1 | (N−1)/N | 0.9688× |
| broadcast | — | 1 | 1.0000× |
Measured on 4×8 fpt-h200 with MeshyLearning.utils.nccl_bench, 32 MiB
all-reduce: 365 µs mean → alg_bw 91.8 GB/s →
bus_bw 91.8 × 1.9375 = 177.9 GB/s. The same run pinned to
NCCL_ALGO=Ring instead of the auto-selected NVLS_TREE: 113.0 GB/s.