NCCL · 4 ranks · buffer split into 4 chunks

Ring Collectives

Each box holds one chunk, and the number inside is its value. Ranks pass chunks to their right-hand neighbour, N−1 times. Watch the numbers: in a reduction they add up, in a gather they only change hands — and the box never changes size either way.

value from rank 0 (10-13)
rank 1 (20-23)
rank 2 (30-33)
rank 3 (40-43)

The dots say whose numbers have been added in — four dots means the chunk is fully reduced. They are not a size: 10 + 20 + 30 + 40 = 100 is still one number occupying one box.

Chunks each rank has sent

Counted in multiples of the whole buffer. One hop moves one chunk = 1/N of the buffer.

Where the factor comes from

Reading the grid

What each one leaves behind

all_reduce
Reduce-scatter, then all-gather. Every rank ends holding every summed chunk — the grid fills with 100, 104, 108, 112. Two passes of N−1 hops, hence 2(N−1)/N. This is the DDP gradient sync.
reduce_scatter
Only the first pass. Each summed chunk lands on exactly one rank, so the output is 1/N the size of the input. FSDP and ZeRO use this: a rank owns one shard of the parameters, so it only needs the summed gradient for that shard.
all_gather
The second pass alone, and the exact inverse of reduce-scatter. Every rank starts holding only its own chunk (the other boxes are empty) and ends holding all four. Nothing is added — the numbers you see are the ones that started elsewhere.
all_to_all
A transpose, not a reduction. Rank i hands its chunk j to rank j, which files it in slot i. No arithmetic at all; what changes is who holds what. This is MoE expert routing, and resharding a tensor from one split axis to another.

Why the factor exists

alg_bw = size / time answers "how fast did my call finish", but it can't be compared across collectives — the same buffer puts different amounts of traffic on the wire. Multiplying by the factor gives bus_bw, the traffic the link actually carried, which is comparable both between collectives and against the hardware's rated bandwidth.

CollectiveHops per rankFactorAt N = 32
all_reduce2(N−1)2(N−1)/N1.9375×
reduce_scatterN−1(N−1)/N0.9688×
all_gatherN−1(N−1)/N0.9688×
all_to_allN−1(N−1)/N0.9688×
broadcast—11.0000×

Measured on 4×8 fpt-h200 with MeshyLearning.utils.nccl_bench, 32 MiB all-reduce: 365 µs mean → alg_bw 91.8 GB/s → bus_bw 91.8 × 1.9375 = 177.9 GB/s. The same run pinned to NCCL_ALGO=Ring instead of the auto-selected NVLS_TREE: 113.0 GB/s.