| --- |
| license: apache-2.0 |
| library_name: stemma |
| pretty_name: Stemma direction model |
| tags: |
| - model-provenance |
| - lineage |
| - safetensors |
| - ai-bom |
| - model-merging |
| - supply-chain |
| - not-a-language-model |
| --- |
| |
| # Stemma direction model |
|
|
| This repository does **not** contain a language model. It contains the small fitted artifacts |
| that the [Stemma](https://github.com/<user>/stemma) provenance tool loads at runtime: |
|
|
| | File | What it is | |
| |---|---| |
| | `direction_model.json` | The fitted `DirectionModel`: `weights` over `stemma.types.DIRECTION_FEATURES`, `bias`, `feature_names`, `scaler_mean`, `scaler_scale`. A regularised logistic combiner — a few dozen floats. | |
| | `sketch_config.json` | The frozen sketch coordinate system: `version` (`SKETCH_VERSION`), `ROLES`, `DEPTH_BUCKETS`, `FEATURES_PER_SLOT`, `SKETCH_DIM`. A sketch computed under a different config is not comparable. | |
| | `sketch_index.npz` + `sketch_index.json` | A prebuilt `stemma.phylogeny.SketchIndex` over the benchmark universe: L2-normalisable float32 sketch vectors, presence masks, model ids and per-model metadata. | |
| | `fit_report.json` | Train/test sizes, held-out accuracy, per-relation accuracy and abstention rate from the fit that produced `direction_model.json`. | |
|
|
| Everything here is generated by `python scripts/push_model.py --repo-id <org>/<name> --bench-dir bench_models --fit`. |
|
|
| ## Intended use |
|
|
| **In scope.** Loading into Stemma to (1) orient a candidate parent/child pair, (2) retrieve |
| candidate parents for a model cheaply from a prebuilt index, (3) produce an AI Bill of Materials |
| that a human then reviews. |
|
|
| ```python |
| from stemma.direction import DirectionModel, estimate_direction |
| from stemma.phylogeny import SketchIndex |
| |
| model = DirectionModel.load("direction_model.json") |
| index = SketchIndex.load("sketch_index") |
| verdict = estimate_direction("org/a", "org/b", model=model) |
| ``` |
|
|
| **Out of scope.** Any automated enforcement, takedown, publication-blocking or procurement |
| decision. Any use as a legal determination of license compliance or infringement. Any claim that a |
| model "is" a derivative of another — Stemma reports how consistent the weights are with a |
| direction of derivation, at a stated confidence, from a small sample of tensors. |
|
|
| ## Generation procedure |
|
|
| 1. **Benchmark construction.** `scripts/build_bench.py` builds real safetensors checkpoints with |
| known lineages — fine-tunes, quantisations, prunings, vocabulary extensions, linear/TIES/DARE |
| merges with known mixing ratios, plus unrelated negatives — and writes `ground_truth.json` with |
| `models`, `edges` and labelled ordered `pairs`. |
| 2. **Feature extraction.** For each labelled ordered pair `(a, b)`, |
| `direction.collect_pair_evidence` Range-reads a handful of shared tensors and |
| `direction.direction_features` reduces them to the antisymmetric feature vector named by |
| `stemma.types.DIRECTION_FEATURES`. The features are antisymmetric by construction: |
| `f(b, a) == -f(a, b)` to within 1e-6, so the fitted combiner cannot learn a positional bias. |
| 3. **Fitting.** `DirectionModel.fit(X, y, l2=...)` on a seeded, deterministic train/test split |
| (default seed 0, 25% held out). `y = +1` when `a` is the parent. Because the features are |
| antisymmetric, each pair is also usable in its mirrored form; the split is done over *pairs*, |
| not over rows, so a pair and its mirror never straddle the split. |
| 4. **Index build.** Every model in the benchmark universe is sketched once |
| (`sketch.sketch_model`) and the resulting vectors are stored in a `SketchIndex`. |
| 5. **Packaging.** `scripts/push_model.py` writes the four files above plus this card and uploads |
| with `HfApi.create_repo(exist_ok=True)` + `upload_folder`. The script is **dry-run by default**: |
| without `--push` it prints exactly what would be uploaded and uploads nothing. |
|
|
| Determinism: every randomised step takes an explicit `seed` (default `0`). |
|
|
| ## Evaluation |
|
|
| Held-out accuracy, per-relation accuracy and abstention rate are written to `fit_report.json` at |
| fit time and mirrored into the repo's README frontmatter-free body by `scripts/push_model.py`. The |
| full benchmark — relatedness AUC and FPR@95TPR, direction accuracy against the symmetric |
| baselines, merge parent-set F1 and mixing MAE, bytes- and seconds-per-decision — is produced by |
| `python benchmarks/run.py` and lives in `benchmarks/results.json`. |
|
|
| No evaluation numbers are quoted in this card. Numbers belong in the generated `fit_report.json` |
| and `benchmarks/results.json`, so that nothing here can drift away from what was actually measured. |
|
|
| Reporting rules the harness enforces (from `docs/FINDINGS.md`): accuracy is reported **per relation |
| type**, abstention is reported alongside accuracy-on-non-abstained, cross-architecture distillation |
| is scored rather than excluded, and the symmetric baselines' 50% direction score is labelled a |
| structural ceiling rather than a tuning failure. |
|
|
| ## Limitations |
|
|
| These are measured, not hypothetical; see `docs/FINDINGS.md`. |
|
|
| - **Direction is near-deterministic only for lossy operations.** Quantisation and pruning scars and |
| vocabulary extension are strong because they are irreversible. Scar-free SFT / LoRA / continued |
| pretraining is **weakly identifiable from two models alone** and relies on outgroup rooting, |
| which needs a usable third relative in the candidate universe. |
| - **Norm growth is recipe-dependent and sign-flips across families.** Measured: |
| `log‖B‖_F − log‖A‖_F` = −0.0171 (**0/8** tensors positive) for `Qwen2.5-0.5B → -Instruct`, and |
| +0.0113 (**8/8** positive) for `SmolLM2-135M → -Instruct`, though both pairs are unambiguously |
| base → instruct-tuned. `norm_growth_asym` is therefore a *fitted* feature with a small weight and |
| never a hand-set sign. |
| - **The Δ-subspace signal is small (~10⁻³–10⁻²) and signed opposite to the naive intuition.** It is |
| a tiebreaker, not a decisive feature. |
| - **Cross-architecture distillation leaves little weight-level signal.** |
| - **TIES / DARE are non-linear**, so mixing ratios recovered by the linear decomposer are |
| approximate; the residual is reported so the mismatch is visible. |
| - **Fitted on a benchmark, not on the wild.** The coefficients reflect the operations and model |
| families `build_bench.py` produced. Expect degradation on architectures, quantisation formats or |
| merge recipes outside that distribution, and refit rather than assuming transfer. |
| - **Sampled evidence.** Sketches read at most two tensors per (role, depth) slot with sub-sampled |
| rows. An adversary who knows the sampling scheme can evade it. This is an auditing aid, not a |
| tamper-proof watermark. |
| - **Version-locked.** A `direction_model.json` is only valid against the `DIRECTION_FEATURES` order |
| it was fitted with, and a `SketchIndex` only against its `SKETCH_VERSION`. Both are recorded in |
| the artifacts and checked on load. |
|
|
| ## Ethics and framing |
|
|
| Stemma reports **statistical evidence** about weight-level similarity and derivation direction, |
| with a confidence attached to every edge. It does **not** establish provenance as fact and does |
| **not** constitute a legal determination of license compliance or infringement. A human must |
| review every finding before any action is taken. Abstention (`direction="unknown"`) is a correct |
| and expected outcome; a confident wrong answer is worse than no answer here. |
|
|
| ## License |
|
|
| Apache-2.0. The artifacts are derived from statistics computed over third-party checkpoints; |
| Stemma does not redistribute any third-party weights, and those checkpoints' own licenses continue |
| to govern them. |
|
|
| --- |
|
|
| ## This build |
|
|
| - Repository: `NagaYu/stemma-direction` |
| - Generated (UTC): `2026-08-08T13:24:27Z` |
| - Sketch format: `stemma-sketch-v1` (dim 1456) |
| - Direction weights: **hand-set priors** (`DirectionModel.default()`), **not fitted**, |
| and that is a deliberate, measured choice rather than a missing step. |
|
|
| ### Why the priors and not a fit |
|
|
| Fitting an L2 logistic combiner on the benchmark's labelled ordered pairs was tried |
| and **lost**. On the same held-out split the priors scored **1.000** accuracy on |
| decided pairs against the fit's **0.500** — chance. Read that with its sample size: |
| the priors abstained on 5 of 7 and decided only 2, so the accuracy gap rests on 2 |
| decisions against 4 and is suggestive, not conclusive. |
|
|
| The decisive evidence is *what the fit learned*. With 13 features and 21 training |
| pairs the problem is underdetermined, and the fit assigned `lattice_asym` a |
| **negative** weight — asserting that the quantised model is the parent. That is |
| physically impossible: dequantisation cannot restore what rounding destroyed, so the |
| scar can only ever appear downstream. It also put its largest weight on the statistic |
| already measured as the weakest. A prior encoding a physical impossibility beats a |
| coefficient fitted on 21 examples. |
|
|
| `--fit` remains available for anyone with a substantially larger labelled corpus: |
| `python scripts/push_model.py --repo-id <id> --bench-dir bench_models --fit` |
|
|
| - Prebuilt index: 20 sketches (backend `faiss`, 0 model(s) unreadable) |
|
|
| Full numbers, including the exact split and per-relation breakdown, are in `fit_report.json` in this repository. The wider benchmark (relatedness AUC / FPR@95TPR, merge F1 and mixing MAE, bytes per decision) is regenerated with `python benchmarks/run.py`. |
|
|
| ## What weight geometry cannot do |
|
|
| Two structural limits were measured after this repository was first published, and they |
| bound how the artifacts here should be used: |
|
|
| 1. **Direction is near-deterministic only for *lossy* operations.** Quantisation, |
| pruning and vocabulary extension score 100%; scar-free SFT/LoRA/CPT edges abstain |
| (mean |llr| ~0.02). The estimator declines rather than guessing, which is the correct |
| failure mode for provenance. |
| 2. **Outgroup rooting is invalid for merge children.** Rooting assumes descendants drift |
| monotonically away from the root, but merging is a *contraction toward the centroid*: |
| `0.6*sft + 0.4*cpt` partly cancels two perturbations and lands **closer to the root |
| than either parent** (root→sft 0.000820, root→cpt 0.001610, root→merge 0.000678). |
| Every correctly chosen sibling outgroup then pushes the answer the *wrong* way. |
| Direction for a merged model must come from the **decomposition**, not from distance |
| geometry — merge precision **1.000**, DARE mixing MAE **0.0004**. |
|
|
| Full derivations, with the measurements that produced them, are in |
| [`docs/FINDINGS.md`](https://github.com/NagaYu/stemma/blob/main/docs/FINDINGS.md). |
|
|
| ## Scope and ethics |
|
|
| Stemma reports **statistical evidence with a confidence**, never a determination of |
| infringement or licence non-compliance. Weight-level similarity and derivation direction |
| are inferences from a small sample of tensors and can be wrong. A human must review every |
| finding before any action is taken. |
|
|