Stemma direction model

This repository does not contain a language model. It contains the small fitted artifacts that the Stemma provenance tool loads at runtime:

File What it is
direction_model.json The fitted DirectionModel: weights over stemma.types.DIRECTION_FEATURES, bias, feature_names, scaler_mean, scaler_scale. A regularised logistic combiner — a few dozen floats.
sketch_config.json The frozen sketch coordinate system: version (SKETCH_VERSION), ROLES, DEPTH_BUCKETS, FEATURES_PER_SLOT, SKETCH_DIM. A sketch computed under a different config is not comparable.
sketch_index.npz + sketch_index.json A prebuilt stemma.phylogeny.SketchIndex over the benchmark universe: L2-normalisable float32 sketch vectors, presence masks, model ids and per-model metadata.
fit_report.json Train/test sizes, held-out accuracy, per-relation accuracy and abstention rate from the fit that produced direction_model.json.

Everything here is generated by python scripts/push_model.py --repo-id <org>/<name> --bench-dir bench_models --fit.

Intended use

In scope. Loading into Stemma to (1) orient a candidate parent/child pair, (2) retrieve candidate parents for a model cheaply from a prebuilt index, (3) produce an AI Bill of Materials that a human then reviews.

from stemma.direction import DirectionModel, estimate_direction
from stemma.phylogeny import SketchIndex

model  = DirectionModel.load("direction_model.json")
index  = SketchIndex.load("sketch_index")
verdict = estimate_direction("org/a", "org/b", model=model)

Out of scope. Any automated enforcement, takedown, publication-blocking or procurement decision. Any use as a legal determination of license compliance or infringement. Any claim that a model "is" a derivative of another — Stemma reports how consistent the weights are with a direction of derivation, at a stated confidence, from a small sample of tensors.

Generation procedure

  1. Benchmark construction. scripts/build_bench.py builds real safetensors checkpoints with known lineages — fine-tunes, quantisations, prunings, vocabulary extensions, linear/TIES/DARE merges with known mixing ratios, plus unrelated negatives — and writes ground_truth.json with models, edges and labelled ordered pairs.
  2. Feature extraction. For each labelled ordered pair (a, b), direction.collect_pair_evidence Range-reads a handful of shared tensors and direction.direction_features reduces them to the antisymmetric feature vector named by stemma.types.DIRECTION_FEATURES. The features are antisymmetric by construction: f(b, a) == -f(a, b) to within 1e-6, so the fitted combiner cannot learn a positional bias.
  3. Fitting. DirectionModel.fit(X, y, l2=...) on a seeded, deterministic train/test split (default seed 0, 25% held out). y = +1 when a is the parent. Because the features are antisymmetric, each pair is also usable in its mirrored form; the split is done over pairs, not over rows, so a pair and its mirror never straddle the split.
  4. Index build. Every model in the benchmark universe is sketched once (sketch.sketch_model) and the resulting vectors are stored in a SketchIndex.
  5. Packaging. scripts/push_model.py writes the four files above plus this card and uploads with HfApi.create_repo(exist_ok=True) + upload_folder. The script is dry-run by default: without --push it prints exactly what would be uploaded and uploads nothing.

Determinism: every randomised step takes an explicit seed (default 0).

Evaluation

Held-out accuracy, per-relation accuracy and abstention rate are written to fit_report.json at fit time and mirrored into the repo's README frontmatter-free body by scripts/push_model.py. The full benchmark — relatedness AUC and FPR@95TPR, direction accuracy against the symmetric baselines, merge parent-set F1 and mixing MAE, bytes- and seconds-per-decision — is produced by python benchmarks/run.py and lives in benchmarks/results.json.

No evaluation numbers are quoted in this card. Numbers belong in the generated fit_report.json and benchmarks/results.json, so that nothing here can drift away from what was actually measured.

Reporting rules the harness enforces (from docs/FINDINGS.md): accuracy is reported per relation type, abstention is reported alongside accuracy-on-non-abstained, cross-architecture distillation is scored rather than excluded, and the symmetric baselines' 50% direction score is labelled a structural ceiling rather than a tuning failure.

Limitations

These are measured, not hypothetical; see docs/FINDINGS.md.

  • Direction is near-deterministic only for lossy operations. Quantisation and pruning scars and vocabulary extension are strong because they are irreversible. Scar-free SFT / LoRA / continued pretraining is weakly identifiable from two models alone and relies on outgroup rooting, which needs a usable third relative in the candidate universe.
  • Norm growth is recipe-dependent and sign-flips across families. Measured: log‖B‖_F − log‖A‖_F = −0.0171 (0/8 tensors positive) for Qwen2.5-0.5B → -Instruct, and +0.0113 (8/8 positive) for SmolLM2-135M → -Instruct, though both pairs are unambiguously base → instruct-tuned. norm_growth_asym is therefore a fitted feature with a small weight and never a hand-set sign.
  • The Δ-subspace signal is small (~10⁻³–10⁻²) and signed opposite to the naive intuition. It is a tiebreaker, not a decisive feature.
  • Cross-architecture distillation leaves little weight-level signal.
  • TIES / DARE are non-linear, so mixing ratios recovered by the linear decomposer are approximate; the residual is reported so the mismatch is visible.
  • Fitted on a benchmark, not on the wild. The coefficients reflect the operations and model families build_bench.py produced. Expect degradation on architectures, quantisation formats or merge recipes outside that distribution, and refit rather than assuming transfer.
  • Sampled evidence. Sketches read at most two tensors per (role, depth) slot with sub-sampled rows. An adversary who knows the sampling scheme can evade it. This is an auditing aid, not a tamper-proof watermark.
  • Version-locked. A direction_model.json is only valid against the DIRECTION_FEATURES order it was fitted with, and a SketchIndex only against its SKETCH_VERSION. Both are recorded in the artifacts and checked on load.

Ethics and framing

Stemma reports statistical evidence about weight-level similarity and derivation direction, with a confidence attached to every edge. It does not establish provenance as fact and does not constitute a legal determination of license compliance or infringement. A human must review every finding before any action is taken. Abstention (direction="unknown") is a correct and expected outcome; a confident wrong answer is worse than no answer here.

License

Apache-2.0. The artifacts are derived from statistics computed over third-party checkpoints; Stemma does not redistribute any third-party weights, and those checkpoints' own licenses continue to govern them.


This build

  • Repository: NagaYu/stemma-direction
  • Generated (UTC): 2026-08-08T13:24:27Z
  • Sketch format: stemma-sketch-v1 (dim 1456)
  • Direction weights: hand-set priors (DirectionModel.default()), not fitted, and that is a deliberate, measured choice rather than a missing step.

Why the priors and not a fit

Fitting an L2 logistic combiner on the benchmark's labelled ordered pairs was tried and lost. On the same held-out split the priors scored 1.000 accuracy on decided pairs against the fit's 0.500 — chance. Read that with its sample size: the priors abstained on 5 of 7 and decided only 2, so the accuracy gap rests on 2 decisions against 4 and is suggestive, not conclusive.

The decisive evidence is what the fit learned. With 13 features and 21 training pairs the problem is underdetermined, and the fit assigned lattice_asym a negative weight — asserting that the quantised model is the parent. That is physically impossible: dequantisation cannot restore what rounding destroyed, so the scar can only ever appear downstream. It also put its largest weight on the statistic already measured as the weakest. A prior encoding a physical impossibility beats a coefficient fitted on 21 examples.

--fit remains available for anyone with a substantially larger labelled corpus: python scripts/push_model.py --repo-id <id> --bench-dir bench_models --fit

  • Prebuilt index: 20 sketches (backend faiss, 0 model(s) unreadable)

Full numbers, including the exact split and per-relation breakdown, are in fit_report.json in this repository. The wider benchmark (relatedness AUC / FPR@95TPR, merge F1 and mixing MAE, bytes per decision) is regenerated with python benchmarks/run.py.

What weight geometry cannot do

Two structural limits were measured after this repository was first published, and they bound how the artifacts here should be used:

  1. Direction is near-deterministic only for lossy operations. Quantisation, pruning and vocabulary extension score 100%; scar-free SFT/LoRA/CPT edges abstain (mean |llr| ~0.02). The estimator declines rather than guessing, which is the correct failure mode for provenance.
  2. Outgroup rooting is invalid for merge children. Rooting assumes descendants drift monotonically away from the root, but merging is a contraction toward the centroid: 0.6*sft + 0.4*cpt partly cancels two perturbations and lands closer to the root than either parent (root→sft 0.000820, root→cpt 0.001610, root→merge 0.000678). Every correctly chosen sibling outgroup then pushes the answer the wrong way. Direction for a merged model must come from the decomposition, not from distance geometry — merge precision 1.000, DARE mixing MAE 0.0004.

Full derivations, with the measurements that produced them, are in docs/FINDINGS.md.

Scope and ethics

Stemma reports statistical evidence with a confidence, never a determination of infringement or licence non-compliance. Weight-level similarity and derivation direction are inferences from a small sample of tensors and can be wrong. A human must review every finding before any action is taken.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using NagaYu/stemma-direction 1