GNN4Colliders / docs /architecture.md
ho22joshua's picture
feat: add initial root transformer model
8344837
|
Raw
History Blame Contribute Delete
3.92 kB

GNN4Colliders architecture

Current architecture

GNN4Colliders is organized around an architecture-neutral event boundary:

ROOT/Awkward
    ↓
EventSample + EventMetadata
    ↓
shared features and representation adapters
    β”œβ”€β”€ GraphSample -> GraphSampleCache -> GraphBatch
    └── SequenceSample -> transformer/token models
    ↓
model family
    ↓
Task -> Trainer/Predictor -> named outputs
Layer Responsibility
data ROOT/Awkward ingestion, metadata, datasets, folds, batching, and caches
features Collider-object features and derived physics quantities
graphs Topology, edge construction, and graph-specific adapters
models Architecture-specific neural networks
tasks Labels, weights, losses, scores, predictions, and metrics
training Optimizers, lifecycle, reproducibility, checkpointing, and distributed utilities
inference Ordered prediction, evaluation, and named output writing
cli / config Thin semantic entry points and configuration composition

Production code must depend on these package boundaries rather than on historical implementation paths. The canonical ROOT-GNN implementation is gnn4colliders.models.root_gnn.EdgeNetwork with FineTunedEdgeNetwork as its transfer boundary. New architecture families must reuse shared data, feature, task, training, and inference interfaces where their representation permits it.

Data and representation boundaries

EventSample and named EventMetadata(fold, weight, sample_id, extra) are representation-independent. GraphSample and GraphBatch are the current graph representation boundary. TensorGraph provides a native tensor path for ROOT-GNN execution on CPU, CUDA, and MPS; DGL remains an optional graph adapter and cache dependency.

SequenceSample and SequenceBatch are the topology-free representation boundary for transformer experiments. TransformerClassifier consumes padded object-token sequences and returns raw logits while reusing shared task, training, checkpoint, and inference interfaces.

The active collider node schema has seven columns: [pt, eta, phi, energy, btag, charge, node_type]. Graph edges are directed, source-major, fully connected without self-loops except for the one-node case, and carry [deta, dphi, dR] features. These are compatibility contracts captured by deterministic tests and the frozen root-gnn-parity-baseline reference fixture.

Configuration and lifecycle

Experiments use semantic configuration such as model.type: root_gnn; model module paths are not part of the new public configuration contract. The CLI delegates to tested application factories and does not contain model or data processing logic.

Checkpoints carry model/task metadata, lifecycle state, schema versions, and optional RNG state. The compatibility package accepts historical checkpoint prefixes and classifier names as a one-way input adapter. No executable historical model code is required at runtime.

Extending the model families

Additional model families should introduce only their representation-specific boundary and model implementation, for example:

data.EventSample -> features -> SequenceSample -> models.root_transformer

It should include a deterministic fixture, unit tests for the representation, an integration path through the shared task/trainer interfaces, and explicit checkpoint/inference behavior. Shared infrastructure should be generalized only when the second model demonstrates a real common use case.

Frozen baseline

The complete ROOT-GNN parity campaign is recorded by the root-gnn-parity-baseline tag. The historical implementation is no longer part of the active source tree. Reference data and compatibility adapters are kept so existing checkpoints and scientific observations remain usable while development moves to new architectures.