stemma-direction / README.md
NagaYu's picture
Publish Stemma direction model, sketch config and prebuilt index
7cca142 verified
|
Raw
History Blame Contribute Delete
10.9 kB
---
license: apache-2.0
library_name: stemma
pretty_name: Stemma direction model
tags:
- model-provenance
- lineage
- safetensors
- ai-bom
- model-merging
- supply-chain
- not-a-language-model
---
# Stemma direction model
This repository does **not** contain a language model. It contains the small fitted artifacts
that the [Stemma](https://github.com/<user>/stemma) provenance tool loads at runtime:
| File | What it is |
|---|---|
| `direction_model.json` | The fitted `DirectionModel`: `weights` over `stemma.types.DIRECTION_FEATURES`, `bias`, `feature_names`, `scaler_mean`, `scaler_scale`. A regularised logistic combiner — a few dozen floats. |
| `sketch_config.json` | The frozen sketch coordinate system: `version` (`SKETCH_VERSION`), `ROLES`, `DEPTH_BUCKETS`, `FEATURES_PER_SLOT`, `SKETCH_DIM`. A sketch computed under a different config is not comparable. |
| `sketch_index.npz` + `sketch_index.json` | A prebuilt `stemma.phylogeny.SketchIndex` over the benchmark universe: L2-normalisable float32 sketch vectors, presence masks, model ids and per-model metadata. |
| `fit_report.json` | Train/test sizes, held-out accuracy, per-relation accuracy and abstention rate from the fit that produced `direction_model.json`. |
Everything here is generated by `python scripts/push_model.py --repo-id <org>/<name> --bench-dir bench_models --fit`.
## Intended use
**In scope.** Loading into Stemma to (1) orient a candidate parent/child pair, (2) retrieve
candidate parents for a model cheaply from a prebuilt index, (3) produce an AI Bill of Materials
that a human then reviews.
```python
from stemma.direction import DirectionModel, estimate_direction
from stemma.phylogeny import SketchIndex
model = DirectionModel.load("direction_model.json")
index = SketchIndex.load("sketch_index")
verdict = estimate_direction("org/a", "org/b", model=model)
```
**Out of scope.** Any automated enforcement, takedown, publication-blocking or procurement
decision. Any use as a legal determination of license compliance or infringement. Any claim that a
model "is" a derivative of another — Stemma reports how consistent the weights are with a
direction of derivation, at a stated confidence, from a small sample of tensors.
## Generation procedure
1. **Benchmark construction.** `scripts/build_bench.py` builds real safetensors checkpoints with
known lineages — fine-tunes, quantisations, prunings, vocabulary extensions, linear/TIES/DARE
merges with known mixing ratios, plus unrelated negatives — and writes `ground_truth.json` with
`models`, `edges` and labelled ordered `pairs`.
2. **Feature extraction.** For each labelled ordered pair `(a, b)`,
`direction.collect_pair_evidence` Range-reads a handful of shared tensors and
`direction.direction_features` reduces them to the antisymmetric feature vector named by
`stemma.types.DIRECTION_FEATURES`. The features are antisymmetric by construction:
`f(b, a) == -f(a, b)` to within 1e-6, so the fitted combiner cannot learn a positional bias.
3. **Fitting.** `DirectionModel.fit(X, y, l2=...)` on a seeded, deterministic train/test split
(default seed 0, 25% held out). `y = +1` when `a` is the parent. Because the features are
antisymmetric, each pair is also usable in its mirrored form; the split is done over *pairs*,
not over rows, so a pair and its mirror never straddle the split.
4. **Index build.** Every model in the benchmark universe is sketched once
(`sketch.sketch_model`) and the resulting vectors are stored in a `SketchIndex`.
5. **Packaging.** `scripts/push_model.py` writes the four files above plus this card and uploads
with `HfApi.create_repo(exist_ok=True)` + `upload_folder`. The script is **dry-run by default**:
without `--push` it prints exactly what would be uploaded and uploads nothing.
Determinism: every randomised step takes an explicit `seed` (default `0`).
## Evaluation
Held-out accuracy, per-relation accuracy and abstention rate are written to `fit_report.json` at
fit time and mirrored into the repo's README frontmatter-free body by `scripts/push_model.py`. The
full benchmark — relatedness AUC and FPR@95TPR, direction accuracy against the symmetric
baselines, merge parent-set F1 and mixing MAE, bytes- and seconds-per-decision — is produced by
`python benchmarks/run.py` and lives in `benchmarks/results.json`.
No evaluation numbers are quoted in this card. Numbers belong in the generated `fit_report.json`
and `benchmarks/results.json`, so that nothing here can drift away from what was actually measured.
Reporting rules the harness enforces (from `docs/FINDINGS.md`): accuracy is reported **per relation
type**, abstention is reported alongside accuracy-on-non-abstained, cross-architecture distillation
is scored rather than excluded, and the symmetric baselines' 50% direction score is labelled a
structural ceiling rather than a tuning failure.
## Limitations
These are measured, not hypothetical; see `docs/FINDINGS.md`.
- **Direction is near-deterministic only for lossy operations.** Quantisation and pruning scars and
vocabulary extension are strong because they are irreversible. Scar-free SFT / LoRA / continued
pretraining is **weakly identifiable from two models alone** and relies on outgroup rooting,
which needs a usable third relative in the candidate universe.
- **Norm growth is recipe-dependent and sign-flips across families.** Measured:
`log‖B‖_F − log‖A‖_F` = −0.0171 (**0/8** tensors positive) for `Qwen2.5-0.5B → -Instruct`, and
+0.0113 (**8/8** positive) for `SmolLM2-135M → -Instruct`, though both pairs are unambiguously
base → instruct-tuned. `norm_growth_asym` is therefore a *fitted* feature with a small weight and
never a hand-set sign.
- **The Δ-subspace signal is small (~10⁻³–10⁻²) and signed opposite to the naive intuition.** It is
a tiebreaker, not a decisive feature.
- **Cross-architecture distillation leaves little weight-level signal.**
- **TIES / DARE are non-linear**, so mixing ratios recovered by the linear decomposer are
approximate; the residual is reported so the mismatch is visible.
- **Fitted on a benchmark, not on the wild.** The coefficients reflect the operations and model
families `build_bench.py` produced. Expect degradation on architectures, quantisation formats or
merge recipes outside that distribution, and refit rather than assuming transfer.
- **Sampled evidence.** Sketches read at most two tensors per (role, depth) slot with sub-sampled
rows. An adversary who knows the sampling scheme can evade it. This is an auditing aid, not a
tamper-proof watermark.
- **Version-locked.** A `direction_model.json` is only valid against the `DIRECTION_FEATURES` order
it was fitted with, and a `SketchIndex` only against its `SKETCH_VERSION`. Both are recorded in
the artifacts and checked on load.
## Ethics and framing
Stemma reports **statistical evidence** about weight-level similarity and derivation direction,
with a confidence attached to every edge. It does **not** establish provenance as fact and does
**not** constitute a legal determination of license compliance or infringement. A human must
review every finding before any action is taken. Abstention (`direction="unknown"`) is a correct
and expected outcome; a confident wrong answer is worse than no answer here.
## License
Apache-2.0. The artifacts are derived from statistics computed over third-party checkpoints;
Stemma does not redistribute any third-party weights, and those checkpoints' own licenses continue
to govern them.
---
## This build
- Repository: `NagaYu/stemma-direction`
- Generated (UTC): `2026-08-08T13:24:27Z`
- Sketch format: `stemma-sketch-v1` (dim 1456)
- Direction weights: **hand-set priors** (`DirectionModel.default()`), **not fitted**,
and that is a deliberate, measured choice rather than a missing step.
### Why the priors and not a fit
Fitting an L2 logistic combiner on the benchmark's labelled ordered pairs was tried
and **lost**. On the same held-out split the priors scored **1.000** accuracy on
decided pairs against the fit's **0.500** — chance. Read that with its sample size:
the priors abstained on 5 of 7 and decided only 2, so the accuracy gap rests on 2
decisions against 4 and is suggestive, not conclusive.
The decisive evidence is *what the fit learned*. With 13 features and 21 training
pairs the problem is underdetermined, and the fit assigned `lattice_asym` a
**negative** weight — asserting that the quantised model is the parent. That is
physically impossible: dequantisation cannot restore what rounding destroyed, so the
scar can only ever appear downstream. It also put its largest weight on the statistic
already measured as the weakest. A prior encoding a physical impossibility beats a
coefficient fitted on 21 examples.
`--fit` remains available for anyone with a substantially larger labelled corpus:
`python scripts/push_model.py --repo-id <id> --bench-dir bench_models --fit`
- Prebuilt index: 20 sketches (backend `faiss`, 0 model(s) unreadable)
Full numbers, including the exact split and per-relation breakdown, are in `fit_report.json` in this repository. The wider benchmark (relatedness AUC / FPR@95TPR, merge F1 and mixing MAE, bytes per decision) is regenerated with `python benchmarks/run.py`.
## What weight geometry cannot do
Two structural limits were measured after this repository was first published, and they
bound how the artifacts here should be used:
1. **Direction is near-deterministic only for *lossy* operations.** Quantisation,
pruning and vocabulary extension score 100%; scar-free SFT/LoRA/CPT edges abstain
(mean |llr| ~0.02). The estimator declines rather than guessing, which is the correct
failure mode for provenance.
2. **Outgroup rooting is invalid for merge children.** Rooting assumes descendants drift
monotonically away from the root, but merging is a *contraction toward the centroid*:
`0.6*sft + 0.4*cpt` partly cancels two perturbations and lands **closer to the root
than either parent** (root→sft 0.000820, root→cpt 0.001610, root→merge 0.000678).
Every correctly chosen sibling outgroup then pushes the answer the *wrong* way.
Direction for a merged model must come from the **decomposition**, not from distance
geometry — merge precision **1.000**, DARE mixing MAE **0.0004**.
Full derivations, with the measurements that produced them, are in
[`docs/FINDINGS.md`](https://github.com/NagaYu/stemma/blob/main/docs/FINDINGS.md).
## Scope and ethics
Stemma reports **statistical evidence with a confidence**, never a determination of
infringement or licence non-compliance. Weight-level similarity and derivation direction
are inferences from a small sample of tensors and can be wrong. A human must review every
finding before any action is taken.