Step #10,000 — 1B model, 32 nodes, DDP
Phase 1 of 4 · ~1.5s elapsed
Node 0 · local NVMe
Node 1 · local NVMe
···
Node 31
Iteration 1 (8 samples each):
G0
G1
G2
···
G255
→ 2,048 samples, no communication
Iteration 2 (grad accumulation):
256 GPUs × 8 samples = 2,048 more samples
→ grad += new_grad
4,096 samples total — still zero communication
DistributedSampler ensures no overlap within each node
GPU 0
local grad
GPU 1
local grad
···
GPU 255
local grad
Tree allreduce
~8 tree levels via Slingshot
GPU-Direct RDMA
GPU → NIC → GPU
~50ms — every GPU has identical avg gradient
CPU completely out of the critical path
GPU 0
AdamW.step()
GPU 1
AdamW.step()
···
GPU 255
AdamW.step()
No comm
Same avg gradient + same state = identical weights
All 256 copies remain perfectly synchronized
4,194,304 tokens digested in ~3.5 seconds
[step 10000 epoch 0.21] loss=2.54 tflops=27.3 mfu=28.5%
tokens_seen: 41.9B
Then do it again — 23,841 more times
4,096 samples × 1,024 tokens × 47,683 steps = 200B tokens total
~46 hours on 32 nodes with Slingshot interconnect
4,194,304
tokens processed (2 accumulation iterations)