Compute node Trainingforward / backward writes ckpt NVMe SSD/mnt/bb/$USER/ Sidecarpolls every 5 min Lustre (Orion)shared filesystem async cp -r stage-in at job start No direct I/O to Lustre during training Zero FS crashes since switching to NVMe
Stage in: copy latest Lustre checkpoint to NVMe on all nodes via srun cp -r