GNN4Colliders / docs /perlmutter.md
ho22joshua's picture
feat: add restartable sharded graph preparation
2c479af
|
Raw
History Blame Contribute Delete
2.83 kB
# Perlmutter execution
The package does not encode Perlmutter paths, modules, or allocation policy.
Create the uv environment in the project location appropriate for your
account, select a site-compatible GPU/driver environment, and use the same
Hydra configuration as on a workstation.
## Single GPU
```bash
uv run gnn4colliders train \
data.cache.path=/path/to/graphs.pt \
environment=perlmutter \
trainer.device=cuda \
trainer.max_epochs=10
```
The equivalent Slurm wrapper is:
```bash
sbatch scripts/slurm/train_single_gpu.sh \
data.cache.path=/path/to/graphs.pt trainer.max_epochs=10
```
## Single-node DDP
```bash
GPUS_PER_NODE=4 sbatch scripts/slurm/train_multi_gpu.sh \
data.cache.path=/path/to/graphs.pt trainer.max_epochs=10
```
The wrapper uses `torchrun`; `batch_size` and `num_workers` are per GPU.
Effective batch size is `data.batch_size * number_of_processes`, and only rank
0 writes the shared checkpoint/config/prediction artifacts.
## Multi-node DDP
```bash
GPUS_PER_NODE=4 sbatch scripts/slurm/train_multi_node.sh \
data.cache.path=/path/to/graphs.pt trainer.max_epochs=10
```
Use the provided script as a template and adapt only allocation/account
settings required by the site. Evaluation and prediction can use
`scripts/slurm/evaluate.sh`; prediction currently gathers moderate-size
results in memory.
## Checks and common failures
```bash
uv run python -c "import torch, dgl; print(torch.__version__, dgl.__version__, torch.cuda.is_available())"
uv run gnn4colliders --help
```
An unavailable DGL wheel, incompatible driver, missing cache, or mismatched
cache schema should be fixed in the environment/input rather than hidden with
package-level path changes. GPU kernels and distributed execution are not
promised bitwise deterministic.
## Large ROOT graph preparation
Large samples should be prepared as persistent shards. Each Slurm array task
reads one contiguous global event range and writes one `shard-NNNNN.pt` file;
completed shards are validated and reused when an interrupted array is
resubmitted. No full-cache in-memory merge is performed.
Create the range manifest once:
```bash
uv run gnn4colliders prepare-manifest \
--config-name config_tth_cp_even_odd \
--shard-count 256 \
--output-dir outputs/ttH_cp_even_odd/shards
```
Submit the CPU array, limiting concurrent tasks to match the filesystem and
allocation capacity:
```bash
sbatch --array=0-255%32 scripts/slurm/prepare_tth_cp_even_odd.sh
```
The script defaults to 64 process workers per task and 256 shards. Set `SHARD_DIR`,
`SHARD_COUNT`, or `CONFIG` at submission time to use another location or
configuration. The current training loader still expects a single cache, so
the next step after extraction is a streaming sharded loader; do not merge the
full benchmark into one `.pt` file.