Perlmutter execution
The package does not encode Perlmutter paths, modules, or allocation policy. Create the uv environment in the project location appropriate for your account, select a site-compatible GPU/driver environment, and use the same Hydra configuration as on a workstation.
Single GPU
uv run gnn4colliders train \
data.cache.path=/path/to/graphs.pt \
environment=perlmutter \
trainer.device=cuda \
trainer.max_epochs=10
The equivalent Slurm wrapper is:
sbatch scripts/slurm/train_single_gpu.sh \
data.cache.path=/path/to/graphs.pt trainer.max_epochs=10
Single-node DDP
GPUS_PER_NODE=4 sbatch scripts/slurm/train_multi_gpu.sh \
data.cache.path=/path/to/graphs.pt trainer.max_epochs=10
The wrapper uses torchrun; batch_size and num_workers are per GPU.
Effective batch size is data.batch_size * number_of_processes, and only rank
0 writes the shared checkpoint/config/prediction artifacts.
Multi-node DDP
GPUS_PER_NODE=4 sbatch scripts/slurm/train_multi_node.sh \
data.cache.path=/path/to/graphs.pt trainer.max_epochs=10
Use the provided script as a template and adapt only allocation/account
settings required by the site. Evaluation and prediction can use
scripts/slurm/evaluate.sh; prediction currently gathers moderate-size
results in memory.
Checks and common failures
uv run python -c "import torch, dgl; print(torch.__version__, dgl.__version__, torch.cuda.is_available())"
uv run gnn4colliders --help
An unavailable DGL wheel, incompatible driver, missing cache, or mismatched cache schema should be fixed in the environment/input rather than hidden with package-level path changes. GPU kernels and distributed execution are not promised bitwise deterministic.
Large ROOT graph preparation
Large samples should be prepared as persistent shards. Each Slurm array task
reads one contiguous global event range and writes one shard-NNNNN.pt file;
completed shards are validated and reused when an interrupted array is
resubmitted. No full-cache in-memory merge is performed.
Create the range manifest once:
uv run gnn4colliders prepare-manifest \
--config-name config_tth_cp_even_odd \
--shard-count 256 \
--output-dir outputs/ttH_cp_even_odd/shards
Submit the CPU array, limiting concurrent tasks to match the filesystem and allocation capacity:
sbatch --array=0-255%32 scripts/slurm/prepare_tth_cp_even_odd.sh
The script defaults to 64 process workers per task and 256 shards. Set SHARD_DIR,
SHARD_COUNT, or CONFIG at submission time to use another location or
configuration. The current training loader still expects a single cache, so
the next step after extraction is a streaming sharded loader; do not merge the
full benchmark into one .pt file.