Step #10,000 — 1B model, 32 nodes, DDP
Phase 1 of 4 · ~1.5s elapsed
Node 0 · local NVMe Node 1 · local NVMe ··· Node 31 Iteration 1 (8 samples each): G0 G1 G2 ··· G255 → 2,048 samples, no communication Iteration 2 (grad accumulation): 256 GPUs × 8 samples = 2,048 more samples → grad += new_grad 4,096 samples total — still zero communication DistributedSampler ensures no overlap within each node
GPU 0local grad GPU 1local grad ··· GPU 255local grad Tree allreduce~8 tree levels via Slingshot GPU-Direct RDMAGPU → NIC → GPU ~50ms — every GPU has identical avg gradient CPU completely out of the critical path
GPU 0AdamW.step() GPU 1AdamW.step() ··· GPU 255AdamW.step() No comm Same avg gradient + same state = identical weights All 256 copies remain perfectly synchronized 4,194,304 tokens digested in ~3.5 seconds
[step 10000 epoch 0.21] loss=2.54 tflops=27.3 mfu=28.5%tokens_seen: 41.9B Then do it again — 23,841 more times 4,096 samples × 1,024 tokens × 47,683 steps = 200B tokens total ~46 hours on 32 nodes with Slingshot interconnect
4,194,304 tokens processed (2 accumulation iterations)