| # Performance notes | |
| The scripts in [`../benchmarks/`](../benchmarks/) separate setup, warmup, and | |
| steady-state timings and emit JSON lines with execution metadata. Float32 eager | |
| execution remains the correctness reference path. | |
| `fully_connected_edges` has a bounded 32-entry cache keyed by node count, | |
| self-loop policy, and device. It preserves source-major ordering and only | |
| reuses topology indices; event-dependent edge features are always recomputed. | |
| The cache is graph-specific and does not alter scientific behavior. | |
| Training already uses `zero_grad(set_to_none=True)` and inference already uses | |
| `torch.inference_mode()` with detached CPU accumulation. | |
| Mixed precision, `torch.compile`, custom kernels, aggressive worker defaults, | |
| and cache-format replacement were not retained without target-machine | |
| measurements. The main known bottleneck is the quadratic graph workload | |
| `N * (N - 1)` and associated DGL message passing; size-aware batching and | |
| streaming prediction remain follow-up work because they affect ordering or | |
| output semantics. | |