GNN4Colliders / docs /perlmutter.md
ho22joshua's picture
feat: add restartable sharded graph preparation
2c479af
|
Raw
History Blame Contribute Delete
2.83 kB

Perlmutter execution

The package does not encode Perlmutter paths, modules, or allocation policy. Create the uv environment in the project location appropriate for your account, select a site-compatible GPU/driver environment, and use the same Hydra configuration as on a workstation.

Single GPU

uv run gnn4colliders train \
  data.cache.path=/path/to/graphs.pt \
  environment=perlmutter \
  trainer.device=cuda \
  trainer.max_epochs=10

The equivalent Slurm wrapper is:

sbatch scripts/slurm/train_single_gpu.sh \
  data.cache.path=/path/to/graphs.pt trainer.max_epochs=10

Single-node DDP

GPUS_PER_NODE=4 sbatch scripts/slurm/train_multi_gpu.sh \
  data.cache.path=/path/to/graphs.pt trainer.max_epochs=10

The wrapper uses torchrun; batch_size and num_workers are per GPU. Effective batch size is data.batch_size * number_of_processes, and only rank 0 writes the shared checkpoint/config/prediction artifacts.

Multi-node DDP

GPUS_PER_NODE=4 sbatch scripts/slurm/train_multi_node.sh \
  data.cache.path=/path/to/graphs.pt trainer.max_epochs=10

Use the provided script as a template and adapt only allocation/account settings required by the site. Evaluation and prediction can use scripts/slurm/evaluate.sh; prediction currently gathers moderate-size results in memory.

Checks and common failures

uv run python -c "import torch, dgl; print(torch.__version__, dgl.__version__, torch.cuda.is_available())"
uv run gnn4colliders --help

An unavailable DGL wheel, incompatible driver, missing cache, or mismatched cache schema should be fixed in the environment/input rather than hidden with package-level path changes. GPU kernels and distributed execution are not promised bitwise deterministic.

Large ROOT graph preparation

Large samples should be prepared as persistent shards. Each Slurm array task reads one contiguous global event range and writes one shard-NNNNN.pt file; completed shards are validated and reused when an interrupted array is resubmitted. No full-cache in-memory merge is performed.

Create the range manifest once:

uv run gnn4colliders prepare-manifest \
  --config-name config_tth_cp_even_odd \
  --shard-count 256 \
  --output-dir outputs/ttH_cp_even_odd/shards

Submit the CPU array, limiting concurrent tasks to match the filesystem and allocation capacity:

sbatch --array=0-255%32 scripts/slurm/prepare_tth_cp_even_odd.sh

The script defaults to 64 process workers per task and 256 shards. Set SHARD_DIR, SHARD_COUNT, or CONFIG at submission time to use another location or configuration. The current training loader still expects a single cache, so the next step after extraction is a streaming sharded loader; do not merge the full benchmark into one .pt file.