Compute nodes have no internet — these prevent hangs
HF_HUB_OFFLINE=1no config downloads
TRANSFORMERS_OFFLINE=1no model downloads
HF_HOME=/lustre/.../hf_cachepre-cached on login node
WANDB_MODE=offlinelocal logs, sync later
The wandb fix (took a full day to figure out)
WANDB_RUN_ID="mamba-1b-32n-v2"deterministic, not random
WANDB_RESUME=allowappend to existing run
Without fixed run ID: 15 fragmented runs
Loss curve
Job 1
Job 2
Job 3
Job 4
5
6
···
Each job restart = new random run ID = broken curve
With fixed run ID: one continuous curve
Loss curve
mamba-1b-32n-v2 (appends across all jobs)
All data in one run. Clean loss curve.
Sync from login node: wandb sync /path/to/run/dir