Stemma direction model
This repository does not contain a language model. It contains the small fitted artifacts that the Stemma provenance tool loads at runtime:
| File | What it is |
|---|---|
direction_model.json |
The fitted DirectionModel: weights over stemma.types.DIRECTION_FEATURES, bias, feature_names, scaler_mean, scaler_scale. A regularised logistic combiner — a few dozen floats. |
sketch_config.json |
The frozen sketch coordinate system: version (SKETCH_VERSION), ROLES, DEPTH_BUCKETS, FEATURES_PER_SLOT, SKETCH_DIM. A sketch computed under a different config is not comparable. |
sketch_index.npz + sketch_index.json |
A prebuilt stemma.phylogeny.SketchIndex over the benchmark universe: L2-normalisable float32 sketch vectors, presence masks, model ids and per-model metadata. |
fit_report.json |
Train/test sizes, held-out accuracy, per-relation accuracy and abstention rate from the fit that produced direction_model.json. |
Everything here is generated by python scripts/push_model.py --repo-id <org>/<name> --bench-dir bench_models --fit.
Intended use
In scope. Loading into Stemma to (1) orient a candidate parent/child pair, (2) retrieve candidate parents for a model cheaply from a prebuilt index, (3) produce an AI Bill of Materials that a human then reviews.
from stemma.direction import DirectionModel, estimate_direction
from stemma.phylogeny import SketchIndex
model = DirectionModel.load("direction_model.json")
index = SketchIndex.load("sketch_index")
verdict = estimate_direction("org/a", "org/b", model=model)
Out of scope. Any automated enforcement, takedown, publication-blocking or procurement decision. Any use as a legal determination of license compliance or infringement. Any claim that a model "is" a derivative of another — Stemma reports how consistent the weights are with a direction of derivation, at a stated confidence, from a small sample of tensors.
Generation procedure
- Benchmark construction.
scripts/build_bench.pybuilds real safetensors checkpoints with known lineages — fine-tunes, quantisations, prunings, vocabulary extensions, linear/TIES/DARE merges with known mixing ratios, plus unrelated negatives — and writesground_truth.jsonwithmodels,edgesand labelled orderedpairs. - Feature extraction. For each labelled ordered pair
(a, b),direction.collect_pair_evidenceRange-reads a handful of shared tensors anddirection.direction_featuresreduces them to the antisymmetric feature vector named bystemma.types.DIRECTION_FEATURES. The features are antisymmetric by construction:f(b, a) == -f(a, b)to within 1e-6, so the fitted combiner cannot learn a positional bias. - Fitting.
DirectionModel.fit(X, y, l2=...)on a seeded, deterministic train/test split (default seed 0, 25% held out).y = +1whenais the parent. Because the features are antisymmetric, each pair is also usable in its mirrored form; the split is done over pairs, not over rows, so a pair and its mirror never straddle the split. - Index build. Every model in the benchmark universe is sketched once
(
sketch.sketch_model) and the resulting vectors are stored in aSketchIndex. - Packaging.
scripts/push_model.pywrites the four files above plus this card and uploads withHfApi.create_repo(exist_ok=True)+upload_folder. The script is dry-run by default: without--pushit prints exactly what would be uploaded and uploads nothing.
Determinism: every randomised step takes an explicit seed (default 0).
Evaluation
Held-out accuracy, per-relation accuracy and abstention rate are written to fit_report.json at
fit time and mirrored into the repo's README frontmatter-free body by scripts/push_model.py. The
full benchmark — relatedness AUC and FPR@95TPR, direction accuracy against the symmetric
baselines, merge parent-set F1 and mixing MAE, bytes- and seconds-per-decision — is produced by
python benchmarks/run.py and lives in benchmarks/results.json.
No evaluation numbers are quoted in this card. Numbers belong in the generated fit_report.json
and benchmarks/results.json, so that nothing here can drift away from what was actually measured.
Reporting rules the harness enforces (from docs/FINDINGS.md): accuracy is reported per relation
type, abstention is reported alongside accuracy-on-non-abstained, cross-architecture distillation
is scored rather than excluded, and the symmetric baselines' 50% direction score is labelled a
structural ceiling rather than a tuning failure.
Limitations
These are measured, not hypothetical; see docs/FINDINGS.md.
- Direction is near-deterministic only for lossy operations. Quantisation and pruning scars and vocabulary extension are strong because they are irreversible. Scar-free SFT / LoRA / continued pretraining is weakly identifiable from two models alone and relies on outgroup rooting, which needs a usable third relative in the candidate universe.
- Norm growth is recipe-dependent and sign-flips across families. Measured:
log‖B‖_F − log‖A‖_F= −0.0171 (0/8 tensors positive) forQwen2.5-0.5B → -Instruct, and +0.0113 (8/8 positive) forSmolLM2-135M → -Instruct, though both pairs are unambiguously base → instruct-tuned.norm_growth_asymis therefore a fitted feature with a small weight and never a hand-set sign. - The Δ-subspace signal is small (~10⁻³–10⁻²) and signed opposite to the naive intuition. It is a tiebreaker, not a decisive feature.
- Cross-architecture distillation leaves little weight-level signal.
- TIES / DARE are non-linear, so mixing ratios recovered by the linear decomposer are approximate; the residual is reported so the mismatch is visible.
- Fitted on a benchmark, not on the wild. The coefficients reflect the operations and model
families
build_bench.pyproduced. Expect degradation on architectures, quantisation formats or merge recipes outside that distribution, and refit rather than assuming transfer. - Sampled evidence. Sketches read at most two tensors per (role, depth) slot with sub-sampled rows. An adversary who knows the sampling scheme can evade it. This is an auditing aid, not a tamper-proof watermark.
- Version-locked. A
direction_model.jsonis only valid against theDIRECTION_FEATURESorder it was fitted with, and aSketchIndexonly against itsSKETCH_VERSION. Both are recorded in the artifacts and checked on load.
Ethics and framing
Stemma reports statistical evidence about weight-level similarity and derivation direction,
with a confidence attached to every edge. It does not establish provenance as fact and does
not constitute a legal determination of license compliance or infringement. A human must
review every finding before any action is taken. Abstention (direction="unknown") is a correct
and expected outcome; a confident wrong answer is worse than no answer here.
License
Apache-2.0. The artifacts are derived from statistics computed over third-party checkpoints; Stemma does not redistribute any third-party weights, and those checkpoints' own licenses continue to govern them.
This build
- Repository:
NagaYu/stemma-direction - Generated (UTC):
2026-08-08T13:24:27Z - Sketch format:
stemma-sketch-v1(dim 1456) - Direction weights: hand-set priors (
DirectionModel.default()), not fitted, and that is a deliberate, measured choice rather than a missing step.
Why the priors and not a fit
Fitting an L2 logistic combiner on the benchmark's labelled ordered pairs was tried and lost. On the same held-out split the priors scored 1.000 accuracy on decided pairs against the fit's 0.500 — chance. Read that with its sample size: the priors abstained on 5 of 7 and decided only 2, so the accuracy gap rests on 2 decisions against 4 and is suggestive, not conclusive.
The decisive evidence is what the fit learned. With 13 features and 21 training
pairs the problem is underdetermined, and the fit assigned lattice_asym a
negative weight — asserting that the quantised model is the parent. That is
physically impossible: dequantisation cannot restore what rounding destroyed, so the
scar can only ever appear downstream. It also put its largest weight on the statistic
already measured as the weakest. A prior encoding a physical impossibility beats a
coefficient fitted on 21 examples.
--fit remains available for anyone with a substantially larger labelled corpus:
python scripts/push_model.py --repo-id <id> --bench-dir bench_models --fit
- Prebuilt index: 20 sketches (backend
faiss, 0 model(s) unreadable)
Full numbers, including the exact split and per-relation breakdown, are in fit_report.json in this repository. The wider benchmark (relatedness AUC / FPR@95TPR, merge F1 and mixing MAE, bytes per decision) is regenerated with python benchmarks/run.py.
What weight geometry cannot do
Two structural limits were measured after this repository was first published, and they bound how the artifacts here should be used:
- Direction is near-deterministic only for lossy operations. Quantisation, pruning and vocabulary extension score 100%; scar-free SFT/LoRA/CPT edges abstain (mean |llr| ~0.02). The estimator declines rather than guessing, which is the correct failure mode for provenance.
- Outgroup rooting is invalid for merge children. Rooting assumes descendants drift
monotonically away from the root, but merging is a contraction toward the centroid:
0.6*sft + 0.4*cptpartly cancels two perturbations and lands closer to the root than either parent (root→sft 0.000820, root→cpt 0.001610, root→merge 0.000678). Every correctly chosen sibling outgroup then pushes the answer the wrong way. Direction for a merged model must come from the decomposition, not from distance geometry — merge precision 1.000, DARE mixing MAE 0.0004.
Full derivations, with the measurements that produced them, are in
docs/FINDINGS.md.
Scope and ethics
Stemma reports statistical evidence with a confidence, never a determination of infringement or licence non-compliance. Weight-level similarity and derivation direction are inferences from a small sample of tensors and can be wrong. A human must review every finding before any action is taken.