Chinchilla optimal for 1B params: ~20B tokens. This run trains at 5× that — intentional. For smaller models, training well past Chinchilla-optimal continues to improve downstream performance. Real wall-clock with checkpointing and HPC chaos: 30–35 hours.