Compute node
Training
forward / backward
writes ckpt
NVMe SSD
/mnt/bb/$USER/
Sidecar
polls every 5 min
Lustre (Orion)
shared filesystem
async cp -r
stage-in at job start
No direct I/O
to Lustre during training
Zero FS crashes
since switching to NVMe
Stage in: copy latest Lustre checkpoint to NVMe on all nodes via srun cp -r