I stared at the progress bar for a solid thirty seconds before I realized something was off.
I had computed max_steps = 23842. I was sure about this number. I had derived it by hand, triple-checked the token count, verified the dataset size. And yet the progress bar said 47,684 total steps. Exactly double.
My first instinct was a bug. Maybe the dataloader was cycling twice. Maybe I had an off-by-one in the epoch calculation. I opened the logs, scrolled to the first optimizer step, and saw:
[step 1 epoch 0.0000] loss=11.24 tflops=27.1 mfu=28.3
Step 1. Not step 2. The optimizer was counting correctly. The progress bar was counting something else entirely. That’s when I learned the difference between optimizer steps and micro-steps, and it sent me down a rabbit hole into how batch sizes actually work in distributed training.
This is the story of that rabbit hole.
The GPU hierarchy you didn’t expect
The cluster I train on has nodes with four GPU cards each. Sounds simple. Four GPUs per node, multiply by node count, done. Except each GPU card has two independent compute dies. Two separate pools of memory, two separate execution engines, sharing one physical package. PyTorch sees each die as a separate device.
So one node isn’t four GPUs. It’s eight.
1 node = 4 physical GPU cards
1 card = 2 compute dies (each with 64 GB HBM)
1 node = 8 "GPUs" in PyTorch's view
32 nodes = 256 GPUs
This matters because every batch size calculation in distributed training starts with one number: the world size. And the world size is the total number of GPU devices across all nodes. For a 32-node job, that’s 256. Not 128. Getting this wrong means your gradient accumulation math is off by 2x, your global batch is wrong, and your learning rate schedule is silently broken.
I got it wrong on my first run.
Four levels of batch size
This is where it gets interesting. There isn’t one “batch size” in distributed training. There are four, and they form a hierarchy. Confusing them is the single most common source of bugs I’ve encountered.
Level 1: the micro-batch
This is what one GPU processes in a single forward pass. One call to model.forward(), one call to loss.backward(). For my 1B parameter model on this hardware, I can fit 8 samples per forward pass per GPU.
per_device_train_batch_size = 8
That’s it. Eight sequences of 1024 tokens each, packed into a tensor of shape [8, 1024], fed through the model in one shot. The micro-batch is constrained entirely by GPU memory.
Level 2: the per-GPU batch
Each GPU might do multiple forward passes before any gradient synchronization happens. This is gradient accumulation: run forward and backward $k$ times, sum up the gradients locally, then sync. The per-GPU batch is:
\[\text{per\_gpu\_batch} = \text{micro\_batch} \times \text{gradient\_accumulation\_steps}\]If gradient_accumulation_steps = 2, each GPU processes $8 \times 2 = 16$ samples before talking to anyone else.
Level 3: the global batch
The total number of samples consumed across all GPUs in one optimizer step. This is the number that matters for training dynamics: the effective batch size that the optimizer sees.
\[\text{global\_batch} = \text{micro\_batch} \times \text{grad\_accum} \times \text{world\_size}\]For my setup: $8 \times 2 \times 256 = 4096$ samples per optimizer step.
Level 4: global tokens
What the model actually learns from, per step:
\[\text{tokens\_per\_step} = \text{global\_batch} \times \text{seq\_length} = 4096 \times 1024 \approx 4.2\text{M tokens}\]Every optimizer step, across all 256 GPUs working in parallel, the model digests 4.2 million tokens. That’s roughly a thousand novels.
The auto-compute trick
Here’s the design decision that took me a while to appreciate. We fix the global batch at 4096 samples, always. Regardless of whether we’re running on 1 node or 128. The training script auto-computes gradient_accumulation_steps to make it work:
per_step_no_accum = per_device_train_batch_size * world_size
gradient_accumulation_steps = global_batch_samples // per_step_no_accum
Why fix the global batch? Because the global batch determines the learning dynamics. The learning rate, the warmup schedule, the gradient noise scale: all of these are calibrated for a specific effective batch size. Change the batch size and you need to re-tune the learning rate. Keep the batch size constant and you can scale to any number of nodes without touching the hyperparameters.
Let me trace through three scenarios to make this concrete.
32 nodes (256 GPUs)
per_step_no_accum = 8 × 256 = 2048
accum = 4096 / 2048 = 2
Each optimizer step:
256 GPUs × 2 forward passes × 8 samples = 4096 ✓
Two forward passes per GPU. Fast.
1 node (8 GPUs)
per_step_no_accum = 8 × 8 = 64
accum = 4096 / 64 = 64
Each optimizer step:
8 GPUs × 64 forward passes × 8 samples = 4096 ✓
Sixty-four forward passes per GPU. Same global batch, same learning dynamics. Just 32x slower wall-clock per step because each GPU is doing 32x more serial work.
128 nodes (1024 GPUs)
per_step_no_accum = 8 × 1024 = 8192
accum = 4096 / 8192 = 0.5 → rounds to 1 (capped)
actual global_batch = 8 × 1 × 1024 = 8192 samples
At this scale we actually overshoot the target. With accum = 1 (the minimum), the global batch becomes 8192 instead of 4096. This is where you’d halve per_device_train_batch_size to 4, bringing the global batch back to $4 \times 1 \times 1024 = 4096$.
The script handles this automatically. The point is: you think in terms of global batch, and the system figures out the rest.
Optimizer steps vs micro-steps
Now we can solve the mystery from the beginning.
An optimizer step is one complete weight update: accumulate gradients across all forward passes and all GPUs, run the allreduce, call optimizer.step(). This is what matters for training. When people say “we trained for 23,842 steps,” they mean optimizer steps.
A micro-step is one forward pass on one GPU. The HuggingFace Trainer progress bar counts micro-steps by default.
\[\text{micro\_steps} = \text{max\_steps} \times \text{gradient\_accumulation\_steps}\]For my 1B model with accum = 2:
There it is. The progress bar said 47,684 because it was counting forward passes, not optimizer steps. The [step N] log lines counted optimizer steps. Both were correct. They were just counting different things.
This gets even more dramatic with different architectures. I was also training an attention-based model on the same cluster. Same 1B parameters, same global batch, same 23,842 optimizer steps. But attention models on this hardware can only fit 2 samples per micro-batch (they use more memory than Mamba). So:
Attention model:
per_device_batch = 2
accum = 4096 / (2 × 256) = 8
micro_steps = 23842 × 8 = 190,736
Mamba model:
per_device_batch = 8
accum = 4096 / (8 × 256) = 2
micro_steps = 23842 × 2 = 47,684
Same training run in terms of optimization. Identical global batch, identical number of tokens seen. But the attention model’s progress bar shows 190,736 steps while Mamba shows 47,684. If you compared these naively: “Mamba is done after 47K steps but attention needs 190K steps”: you’d draw completely wrong conclusions about convergence.
Always compare optimizer steps. Never micro-steps.
Token accounting
Let’s close with the full accounting. How many tokens does this training run consume?
\[\text{total\_tokens} = \text{max\_steps} \times \text{global\_batch} \times \text{seq\_length}\] \[= 23842 \times 4096 \times 1024 = 100{,}000{,}595{,}968 \approx 100\text{B tokens}\]For context: the Chinchilla scaling law suggests a 1B parameter model should see roughly 20B tokens for compute-optimal training. We’re at 5x that. This is intentional. For smaller models, training well past the Chinchilla-optimal point continues to improve downstream performance. It’s a common practice: train small models on way more data than theory suggests, because the marginal cost of more tokens at 1B scale is small relative to the quality gains.
At 4.2M tokens per step and roughly 3.5 seconds per step, the aggregate throughput is about 1.2 million tokens per second across all 256 GPUs. That’s about 4,700 tokens per second per GPU. The full 100B token run takes around 23 hours on 32 nodes.
100B tokens ÷ 1.2M tokens/sec ≈ 83,333 sec ≈ 23.1 hours
Twenty-three hours. Not bad for 100 billion tokens. But that’s assuming zero overhead from checkpointing, evaluation, node failures, and job restarts. The real wall-clock time is more like 30-35 hours once you account for the chaos of running on shared HPC infrastructure. That chaos is a story for another day.
What’s next
This was Part 1: the batch size hierarchy and the arithmetic that makes distributed training work. The numbers are mechanical once you understand the four levels, but getting them wrong silently corrupts your training.
In Part 2, I’ll cover DDP vs FSDP (how gradients and parameters actually move between GPUs), data sharding to NVMe (why we don’t let 256 GPUs read from the parallel filesystem simultaneously), and the multi-node launch pattern that trips up everyone at least once.