1
Optimizer steps
Training length in weight updates
23,842
×
2
Global batch
Samples per optimizer step
4,096
×
3
Sequence length
Tokens per sample
1,024
=
T
Total tokens
23,842 × 4,096 × 1,024
~100B
4.2M
tokens per step
1.2M/s
aggregate throughput
~23 hrs
ideal wall-clock
Chinchilla optimal for 1B params: ~20B tokens. This run trains at 5× that — intentional. For smaller models, training well past Chinchilla-optimal continues to improve downstream performance. Real wall-clock with checkpointing and HPC chaos: 30–35 hours.