Performance notes
The scripts in ../benchmarks/ separate setup, warmup, and
steady-state timings and emit JSON lines with execution metadata. Float32 eager
execution remains the correctness reference path.
fully_connected_edges has a bounded 32-entry cache keyed by node count,
self-loop policy, and device. It preserves source-major ordering and only
reuses topology indices; event-dependent edge features are always recomputed.
The cache is graph-specific and does not alter scientific behavior.
Training already uses zero_grad(set_to_none=True) and inference already uses
torch.inference_mode() with detached CPU accumulation.
Mixed precision, torch.compile, custom kernels, aggressive worker defaults,
and cache-format replacement were not retained without target-machine
measurements. The main known bottleneck is the quadratic graph workload
N * (N - 1) and associated DGL message passing; size-aware batching and
streaming prediction remain follow-up work because they affect ordering or
output semantics.