| # Performance benchmarks |
|
|
| These scripts use fixed synthetic inputs, explicit seeds, warmup iterations, |
| and JSON-lines output. GPU timings synchronize after every measured operation; |
| warmup and one-time setup costs are excluded from steady-state numbers. |
|
|
| ```bash |
| uv run python benchmarks/benchmark_preprocessing.py |
| uv run python benchmarks/benchmark_dataloader.py |
| uv run python benchmarks/benchmark_training.py --device cpu |
| uv run python benchmarks/benchmark_inference.py --device cpu |
| uv run python benchmarks/benchmark_training.py --device cuda --profile |
| ``` |
|
|
| DGL-dependent scripts report a structured `skipped` result when the optional |
| `root-gnn` extra is absent. Profiler traces go under `profiles/` and are not |
| committed. |
|
|
| ## Baseline measurements |
|
|
| The portable baseline in this checkout uses Python 3.12.13, PyTorch 2.2.2+cu121, |
| DGL 2.4.0+cu121, CPU, `nodes=32`, `iterations=10`, and `warmup=3`. Exact timings are machine |
| dependent; the JSON output from a local run is authoritative. |
|
|
| | Measurement | Mean | Median | |
| | --- | ---: | ---: | |
| | feature construction | 0.47 ms/event | 0.46 ms/event | |
| | edge features | 0.11 ms/graph | 0.11 ms/graph | |
| | graph-sample loader | 1.65 ms/iteration | 1.65 ms/iteration | |
| | ROOT-GNN training step | 21.19 ms/step | 5.32 ms/step | |
| | ROOT-GNN inference | 1.48 ms/graph | 1.49 ms/graph | |
|
|
| On the available NVIDIA A100-PCIE-40GB with CUDA 12.1 runtime, the same small |
| synthetic ROOT-GNN benchmark measured 6.41 ms/step (17.3 MiB peak allocated) |
| and 2.00 ms/graph. These are microbenchmarks, not production-workload claims. |
|
|
| DGL is available in this environment. The training mean is skewed by one CPU |
| warmup-adjacent outlier; median is the more useful steady-state indicator. DDP |
| measurements require a multi-process run and are not inferred from CPU numbers. |
|
|
| The topology cache is bounded to 32 node-count/device/policy entries and only |
| stores reusable index tensors, never event-specific edge features. |
|
|
| On the same CPU, a direct topology microbenchmark for 64 nodes measured |
| 201.6 microseconds per cold construction versus 10.5 microseconds for a warm |
| cache lookup (about 19x for this isolated operation). This is a targeted |
| index-construction result, not an end-to-end training speedup; graph feature |
| construction and DGL message passing remain separate costs. |
|
|