Preemption-safe training template
A minimal PyTorch training loop that survives spot / interruptible GPU preemption. Copy train.py and swap in your model and data.
What it does
- Keeps all recovery state (weights, optimizer, RNG, step) in one directory:
$NODUS_STATE_DIRor./state. - Writes checkpoints atomically (temp file, then
os.replace). - Resumes from the last checkpoint on startup.
- Saves on a timer, and again immediately when the platform sends a checkpoint request.
Run locally
pip install torch
python train.py # Ctrl+C midway, run again: it resumes
Run on a cheap interruptible GPU with Nodus
Nodus copies the state directory off the machine before a reclaim and restores it on the replacement, so preemptions cost minutes instead of the whole run.
pip install nodus-compute
nodus login
nodus run --gpu L4 --image nodus/pytorch --interruptible --checkpoint /nodus/state -d -- python train.py
nodus logs -f job/<name>
New accounts get a $30 starter grant (valid for 30 days), so the first runs need no card. Add --dry-run to see the cost estimate before anything starts.
Picking a checkpoint interval
Aim for at least 4 checkpoints per expected run and keep time spent saving under about 10% of runtime. Set SAVE_EVERY_SEC accordingly.
More
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support