G0 G1 G2 G3 G4 ··· ··· G255 Each GPU sends a chunk to its neighbor, accumulates from the other side Optimal bandwidth O(N) latency: 510 steps
Latency (256 GPUs)
510 steps
Bandwidth efficiency
Optimal
Used at scale?
No — too slow