# Paretrix Quantization Suite — Documentation Technical reference for **Paretrix v2.6.0**: architecture, pipeline, calibration ledger, knowledge base, campaign engine, modules, and the evaluation standard. Benchmarks live in [README.md](README.md), the single canonical measurement archive. --- ## 1. Philosophy & Architecture ### 1.1 The Pareto principle applied to quantization Every quantization decision is a trade: bytes freed against fidelity lost. Paretrix measures that trade per tensor class (ΔKLD/MiB) at a real operating point, then allocates under an exact budget so that **no better trade exists at that size** (the Pareto frontier of the (MiB, KLD) plane). The shipping name states its weight band, and the quality objective optimizes inside that band. Three mechanisms deliver this: 1. **Flat recipes** where uniformity wins (mid-band and top-band): a structural base type plus measured levers (recurrent floors, band raises, readout levers), emitted as `llama-quantize` regex rules. 2. **Exact-budget DP** where heterogeneity pays (low-band and squeeze campaigns): a multiple-choice knapsack solved to the MiB, priced by measured notch rates and measured buy rates. 3. **Duels at the verdict regime** decide adoption: a candidate takes the rung name only on a KLD gain at or above 1×MASD (floor 0.001) at ctx 4096×8. ### 1.2 THE MATRIX — three-level knowledge base | Level | File | Scope | Precedence | | :---: | :--- | :--- | :---: | | 1 | `model-paretrix.json` | This model's measured tables (anchors, probes, rates, allocations, cocktails, champions) | First | | 2 | `arch-paretrix/-paretrix.json` | Per-architecture shared line: banded sell/buy tables, facts, corrections, witnesses | Second | | 3 | `arch-paretrix/generic-paretrix.json` | Universal prior (banded mid/deep, distilled from six campaign families) | Third | Each level overlays the one below **class by class**: a class that an upper level never measured prices from the level below. Shared levels (2, 3) are **size-normalized**: `rate × line_reference_mib / model_mib`. A 1.5B and a 9B model of the same architecture therefore share one line. Measured effect: the between-checkpoint rate spread tightens from ×4.7 to ×2.1. Lines are **banded**: mid (floors at or above Q4_K) and deep (IQ4_XS and below), selected per group by its floor depth. ### 1.3 Rate-exchange theory At any operating point, every class carries two measured prices: sell (KLD/MiB freed by a notch cut) and buy (KLD/MiB gained by a raise). Allocation quality depends on that exchange. With both ends noise-clamped (free-sell and dead-buy clamps), the exchange is **win-or-tie by construction** against its own anchor: ties mark inversion-free basins (T1). Anchor choice dominates the table you see (T2); band rates are ctx-dependent on SWA topologies (T3); fine-tunes flatten the field (T4). The default path trades on priors and the campaign trades on measurement, and the D-gap measures the prior's error at the operating point (T5). --- ## 2. Pipeline Overview ``` [HuggingFace checkpoint (safetensors + sidecars)] │ ▼ 01_SAFETENSORS-to-BF16-GGUF.py [/Paretrix/model-BF16.gguf] + [mtp-BF16 · mmproj-BF16 · dspark/dflash-BF16] │ ▼ 02_BF16-GGUF-to-Q8-imatrix.py [model-Q8_0.gguf] + [imatrix.gguf] │ ├─▶ 03_BF16-GGUF-to-Paretrix.py → [model-Paretrix--pc.gguf] │ [+ classic Q8_0 / Q6_K / Q5_K_M / IQ4_XS / IQ3_M] ├─▶ Paretrix-modules.py → [mtp / mmproj / dspark / dflash-Paretrix-.gguf] └─▶ 14_gguf-module-fusion.py → [fused deployment GGUF] │ ▼ 11_perplexity-test.py (verdict regime 4096×8) [PPL · KLD · RMS Δp · top-p · Pareto frontier] │ ▼ Paretrix.py duel --adopt · Paretrix.py receipt · Paretrix.py matrix [model-paretrix.json] → [arch-paretrix/-paretrix.json] → next model ``` Orchestrated by `00_Paretrix-pipeline.py`, which resumes from its working folder (the folder holds the state). --- ## 3. Script Reference ### 3.1 `00_Paretrix-pipeline.py`: orchestrator and shared library Runs 01 → 02 → 03 with resume semantics: existing outputs skip automatically. The script also serves as **PARETRIX COMMON**, the shared library that every other script imports as `PC` (console grammar, live-line child runner, GGUF helpers, file utilities, and the re-shuffle composer). Module repos join the same run: 01 converts each extra folder into the working root, and the module is quantized immediately. ```bash python 00_Paretrix-pipeline.py [module-repo ...] python 00_Paretrix-pipeline.py MyModel MyModel-DSpark --all python 00_Paretrix-pipeline.py MyModel/Paretrix --from 03 python 00_Paretrix-pipeline.py MyModel --dry-run ``` Key flags: `--from {01,02,03}` · `--all` · `--tier a,b` · `--paretrix-only` · `--no-campaign` · `--no-audit` · `--no-duel` · `--no-cocktail` · `--no-fold` · `--no-modules` · `--module-profile {balanced,compact}` · `--outtype {bf16,f16}` · `--cpu-only` · `--force` · `--keep-going` · `--dry-run`. ### 3.2 `01_SAFETENSORS-to-BF16-GGUF.py`: conversion Converts safetensors into a pristine BF16 GGUF with a **native split**: | Output | Content | | :--- | :--- | | `model-BF16.gguf` | Text model; MTP head and vision tower excluded | | `mtp-BF16.gguf` | Standalone MTP fusion head (when present) | | `mmproj-BF16.gguf` | Vision projector (multimodal models) | | `dspark/dflash-BF16.gguf` | Draft repos convert whole, under their own stem | The full model (text + head) is a build intermediate, removed once the head is extracted. `--keep-full` retains it, and 14 rebuilds it exactly. If the converter rejects `--no-mtp`, the full conversion runs first and a structural MTP strip produces the text model. Paretrix accepts raw BF16 sources only; pre-quantized checkpoints (qweight/scales/qzeros keys) are refused. Also handles: RoPE config normalization, remote-code compatibility patching, preprocessor config completion, nanbeige padded-vocab compatibility, shard-index healing, single-file staging, and immediate module quantization through `Paretrix-modules.py` (`--no-modules` defers it). ### 3.3 `02_BF16-GGUF-to-Q8-imatrix.py`: Q8 source and imatrix Produces `model-Q8_0.gguf` (kept for 03's archival classic) and `imatrix.gguf` (activation statistics for the quantization grid). The Q8_0 source speeds up offloadable layers and CPU dot products, while the statistics stay quasi-lossless (measured ρ 0.99999, top-100 overlap 100/100; see L13). `IMATRIX_SOURCE_QUANT=bf16` runs on the BF16 source instead. GPU autotune opens a descending probe ladder at the computed geometry boundary (fit-era slot accounting, calibrated fixed reserve). It then runs iso-chunk recurrent split validation (a broken GDN split multiplies PPL) and verifies PPL on the final run. Every candidate carries an explicit `-ngl` (0 = CPU). Unified CUDA memory stays off, so an out-of-memory condition surfaces as a hard allocation failure. Corpus: `datasets/bartowski-imatrix-v5-semantic.txt` (downloaded once, or adopted from the legacy folder). ### 3.4 `03_BF16-GGUF-to-Paretrix.py`: the Paretrix core Two phases per model: **Phase 1: caches.** L0 architecture audit (10, first contact only) · arch line loaded (or noted absent) · model cache created or loaded. **Phase 2: champions** (`model-Paretrix--pc.gguf`), per armed tier: 1. Champion on disk → skip (with the R19 stock-twin check). 2. Allocation in the model cache → direct build. 3. Fresh champion above (L15) → campaign (preflight → anchor → notch probes → rates → exact-budget DP → build → duel → adopt → cocktail). 4. Neither → engine default path (flat recipes / DP from priors). The script also emits the **classic line** (Q8_0 archival from 02, plus low-band `Q6_K/Q5_K_M/IQ4_XS/IQ3_M`-imx on lever-less topologies), and closes with the arch fold (`Paretrix.py matrix`). Key flags: `--all` · `--tier` · `--paretrix-only` · `--no-campaign` · `--no-audit` · `--no-duel` · `--no-cocktail` · `--no-fold` · `--force` · `--dry-run`. ### 3.5 `Paretrix.py`: engine and campaign CLI The quantization engine (classifier, flat recipes, exact-budget DP, Gumbel-STE search) and the campaign namespace: | Subcommand | Purpose | | :--- | :--- | | `campaign` | Anchor → notch probes → rates → DP → build (one rung, one command) | | `duel` | Measures candidate vs incumbent at the verdict regime; `--adopt` takes the rung name | | `adopt` / `receipt` | Naming and receipt management (`--scan` backfills from sidecars and the eval cache) | | `cocktail` | Composes a re-shuffle from measured arms (overlap exclusion, value-positive bundles) | | `build` | Materializes an allocation sidecar (`--from-cache` for direct rebuilds) | | `verify` | Plan-vs-artifact tensor type comparison | | `polish` | In-format STE polish (measured: 0/25 tensors improved; kept as a falsified arm) | | `matrix` | Folds every model cache into the arch lines (THE MATRIX) | | `cache` / `compact` | Inspects the knowledge base · compresses tensor maps (col1 format, ~20–30× smaller) | | `atlas` | Knowledge base evidence: rate field, lineage, probe economy, routes, lint | | `digest` | Champion composition digest (class × tier × depth per rung) | | `diff` | Byte-level GGUF regression check (tensors + KV) | | `selftest` | DP engine, manifold ops, STE loop | Legacy entry points: `Paretrix.py --model … --profile … --size … --run` (the allocator CLI) and `Paretrix.py --tier all` (the tier runner). ### 3.6 `Paretrix-modules.py`: MTP · DFlash · DSpark · MMProj Imatrix-free module quantization. Kinds are auto-detected from the GGUF (`--kind` overrides): `draft` (DFlash/DSpark), `mtp`, and `mmproj`. Role-based floors apply to draft and MTP modules (bridge ≥ Q6_K, proposal ≥ Q5_K, norms F16/F32, 1-D tensors pinned F32, unclassified tensors Q5_K). MMProj follows a profile grid (critical tensors pinned F32, deep-boost on trailing blocks). Available types are K-quants, Q8_0, and F16/F32 via gguf-py, with a native `llama-quantize` fallback and artifact-vs-plan verification. ```bash python Paretrix-modules.py --model mtp-BF16.gguf python Paretrix-modules.py --model dspark-BF16.gguf --profile compact python Paretrix-modules.py --model mmproj-BF16.gguf --dry-run ``` ### 3.7 `10_arch-inspect.py`: architecture audit Five lenses on a GGUF or safetensors folder: 1. **Tensor inventory**: class × count × dtype × MiB × shape. 2. **Per-layer signatures**: deviant layers (MTP heads, hybrids). Periodic patterns are read as hybrid design; only true anomalies raise a warning. 3. **Novelty scan**: suffixes outside llama.cpp's known taxonomy. 4. **Binary support cross-check**: arch and every tensor name template checked against `llama-arch.cpp`. 5. **Imatrix coverage**: quantizable weights without activation statistics (vectors and lookup tables count as by-design and are not listed). `--brief` (used by 03) shows the top classes and the verdict lines. ### 3.8 `11_perplexity-test.py`: perplexity and fidelity sweep Sweeps every product in a working folder at the verdict regime: PPL, KLD (vs BF16 reference logits), RMS Δp, top-p, t/s, chunk stability (tail σ / MASD), and the **Pareto frontier column**. It offers an interactive menu (`a` = all, indices, names, globs) or piped stdin for automation (`echo a | python 11_…`). - **KL bases** are per-regime (`kld-bf16-ctx4096.dat`, etc.) and per-model, generated once. - **Eval cache** (`eval-cache.json`): KL results keyed on file identity × regime × KL base × binary build, so duels and sweeps reuse measurements. - **Offload planning** inherits 02's calibrated geometry (fit-era slot accounting, KL-mode reserve, recurrent split validation via 02's iso-chunk protocol). - **Regimes**: default `4096×8` (Long Horizon, L18 verdict) · `--medium` `2048×16` · `--light` `1024×32` · `--short` `512×64` (probe regime). Span = 32,768 tokens across all. ### 3.9 `12_attribution-probe.py`: marginal knockout probes Measures ΔKLD/MiB of single-class levers at an operating point. It quantizes the arm, evaluates at the probe regime (512×64), and records the result in a resumable CSV. Campaign probes run with `--no-summary`; the batch closes with one `--summary-only` attribution table (solo screen, exchange rate, shrink/re-shuffle chains). The inert-arm guard detects type-identical builds. `PARETRIX_PROBE_NGL` controls offload. ### 3.10 `13_draft-acceptance-sweep.py`: A/B acceptance battery Spawns `llama-server` with and without the module, replays fixed prompts at temperature 0, and records acceptance, mean length, and throughput per `--n-max` into `draft-acceptance.csv`. The trained block is read from the module metadata (`*.block_size`). `--chat` wraps probes in the target chat template. `--ngl` is recorded, and baselines are compared only at matching offload (L20). The thermal guard flags baselines that drift from the reference. ### 3.11 `14_gguf-module-fusion.py`: module fusion Merges 2+ GGUF files into one. The first file wins on duplicates, and `block_count` plus per-layer arrays widen when a merged head extends the layer index range, touching `.block_count` only, never a sub-module's count. MTP folder mode rebuilds the `model-mtp-*` deployment files. Output lands as `.part` and renames on completion. Float outputs receive Xpress8K NTFS compression. ```bash python 14_gguf-module-fusion.py --mtp [--target ] [--head ] python 14_gguf-module-fusion.py out.gguf main.gguf extra.gguf [...] ``` --- ## 4. The Mod-3 Ladder Ratios are multiples of 3 percent of BF16, listed from largest to smallest. The shipping name carries the ratio (`-pc`), enforced at ±1.5 pp (R14). A build outside every band is withdrawn. A build inside a neighboring band re-homes to that rung, or is withdrawn if the target name is already held, so a rung name never ships outside its band. The gaps (37.5–40.5, 43.5–46.5) own no rung. | Tier | Ratio | Base | Key levers | | :--- | :---: | :--- | :--- | | Fidelity | 48% | Q6_K | + Q8_0 pockets biased to the late layer third (G4) | | Precision | 42% | Q6_K | + band Q8_0; readout lever arch-scoped (L12) | | Quality | 36% | Q5_K_M | L6 flat + imatrix + recurrent floors (inert under regex rules on stateless topologies) | | Compact | 33% | Q4_K_M | Flat + band Q6_K + gate Q6_K + recurrent floors, or DP | | Mini | 30% | IQ4_XS | L4 flat on GDN (ssm_out Q6_K, α/β Q8_0, band Q6_K) or DP | | Nano | 27% | IQ3_S–IQ4_XS | DP with measured rates | | Pico | 24% | IQ3_S | Band Q5_K + gate Q6_K (flat) or DP | | Femto | 21% | IQ3_XXS | Deep-squeeze DP; campaign product | Environment knobs per rung: `PARETRIX_INCLUDE_PICO/FEMTO` (scope: state-rich ≥3B) · `PARETRIX_INCLUDE_PRECISION/FIDELITY` (default on) · `PARETRIX_QUALITY_ENGINE` (`native` flat / `allocate` DP) · `PARETRIX_FEMTO_CLASSIC` · `PARETRIX_FEMTO_RECIPE`. --- ## 5. Calibration Ledger Binding laws (L) and rules (R), measured across nine families (state-space hybrids, dense SWA networks, looped untied trunks, shortconv mixers, fine-tunes, and vision-aligned backbones). The code ledger in `Paretrix.py` is the living source; this section is its reference form. ### 5.1 Laws - **L0: Pre-adoption audit.** Untied lm_heads expose `output.weight`. Run 10 before conversion so the readout is classified correctly instead of parking at the worst tier (normalized through the `readout` class since 2.5.0). - **L1: Readout and embedding value.** Keep tied `token_embd` at Q6_K or higher on production tiers: 94 MiB per top-p point (93% of Q8_0's gain). - **L2: FFN information bottleneck.** At ≤33% budget, `ffn_down` is the primary bottleneck. Sub-IQ4_XS cuts cost ~2e-4 KLD/MiB, an order of magnitude worse than any winning promotion. ReaderLM-v2 (75% FFN mass) validated the Q4_K floor restoring monotonicity. - **L3: Depth prior.** Surplus flows to early recurrent state, then to full attention (scope: ≤33%). In the top band, Q8_0 pockets land in the **late** layer third in 6/6 families (G4 refinement). Deep-band ffn_gate_up cuts land early in 4/4. - **L4: Recurrent readout floor.** Keep `ssm_out` at Q6_K or higher on production tiers: sub-Q6_K collapses output entropy (Ornith: KLD 0.081 → 0.0241 once restored). One exception is measured: the Xiaomi cocktail sells `ssm_out` to Q5_K and still wins, so the result holds per family and does not generalize. - **L5: Small-model constructibility.** ≥9B at Mini · 3B–4B at Compact · ~1B at Quality. 0.33–0.40 KLD is the entropy-dissolution ceiling (three convergent families). Restrict `--allow-q3-or-lower` to grids at or above 25%. - **L6: Mid-band uniformity (price-scoped).** Proxy-priced upgrades pay a heterogeneity tax: Quality-36 delegates to flat Q5_K_M. With both sides rate-measured, the DP reclaims the band (lfm2 −13% relative, nanbeige −34% relative). - **L7: Tied readout invariance.** `token_embd` doubling as lm_head rides the output floor (Q6_K). Its lever works as an additional purchase, not as a swap inside a pinned budget. - **L8: Cut locality over depth.** Scarce full-attention bands are the currency below ~30%. Above that, the uniform body dominates. - **L9: Readout economy.** An untied lm_head rides its class floor. Input-side embd purchases above Q5_K price as noise on most families (3–10× below the body step). Exception: nanbeige-4.2 (2.68e-4/MiB, comparable to its body step). - **L10: Unmeasured classes** follow their declared floors, not the upgrade queue. - **L11: Speculative verification bound.** Draft quantization moves acceptance by at most 0.03 (nanbeige +0.0104, lfm2 +0.0055, MiniCPM5 bit-identical) and speeds the draft itself (+13–21% t/s). - **L12: Readout lever.** High-precision Q8_0 on the logits matrix (tied `token_embd`, untied `lm_head`) delivers measurable gains. This is an arch-scoped correction: spark2_5 carries it, and the qwen35 default downgrades it at the top band. - **L13: Rate-calibrated allocation.** Notch probes and measured buys feed the exact-budget DP (budget hit to ±1 MiB). Rates transfer across anchors within ~5% on bulk classes. - **L14: Anchor-basin economics.** A DP squeeze re-ranks within the anchor's basin, so anchor choice dominates (same Mini slot: 0.0687 fresh vs 0.1035 consumed). - **L15: Nearest-fresh-champion targeting.** The anchor is the nearest champion above that has not been through a deep squeeze (flat/allocator lineage). Reach is rate-mediated, not distance-mediated. - **L16: Top-band buy saturation.** Measured buy rates extend linearly only to Q8_0; F16 buys fall back to the MSE proxy. The top band's sells are nearly free and its buys real, so re-shuffles win where buys are cheap. - **L17: Self re-shuffle primacy (conditional).** At a defended rung, attack with the self re-shuffle. The incumbent's builder lacks the RCO vocabulary, and the self anchor is priced at the operating point. This holds on a live buy table. When ≥3 of 5 buys are dead (saturated basin), the fresh champion above wins instead. Rate test (L17-bis): the re-shuffle wins if and only if the max affordable live buy rate exceeds the min sell rate. After R13, it wins or ties by construction. - **L18: SWA long-context binding.** On windowed-attention topologies, the ctx-512 duel misprices band-heavy allocations in both directions (Quality +0.0005@512 → −0.0107@4096; Pico −0.0361@512 → −0.0033@4096). Verdicts bind at the native context, and band rates are regime-bound. - **L19: Fine-tune field flattening.** A deep fine-tune compresses the rate field: sells fall 3–10× under the base priors, and most buys die (NeoHorse 7/7). Campaign value concentrates in deep sheds and self re-shuffles. Distinguish this from dead-lever families: flat sells with live buys point to L12/L16, not L19. - **L20: Draft-target pairing and regime term.** Acceptance is a property of the (draft, target) pair and the offload regime (±0.02, four times the L11 drift). Read acceptance only at matched `ngl`. Pair each draft with the least-quantized, highest-ngl target the budget allows. Draft-side budget: weights + a fixed block graph (`block_size=16` → ≈1 GiB) + ≈0.1 GiB KV. ### 5.2 Rules - **R3: Loop-neutral depth prior.** On looped trunks (`num_loops > 1`), keep the early-layer downgrade prior disabled. The imatrix integrates every visit, so the prior runs against the measured gradient (nanbeige: blk.0 at 4.1% vs blk.21 hottest, a 24.45× gradient). - **R6-bis: Readout-share parity.** When an untied readout pair exceeds 30% of the mass, demote the input side to stock Q4_K parity and reinvest the freed bytes inside the pool. Top-p corollary: cap readout sells at one notch, because the d2 notch causes a top-p cliff that the KLD line does not price. - **R6-ter: Tied vocab share.** A tied embd notch saves 0.066 × share of the model: 2.2 pp for Xiaomi (34%), 1.0 pp for NeoHorse (15%), and 0.5 pp for Spark (8%). At a share ≥ 15%, the floors are Q5_K at Mini and Nano, Q4_K at Pico, and IQ4_NL at Femto (every campaign champion sits there). Below 15%, Q6_K holds (at 4096 ctx, Spark's flat Pico won). `PARETRIX_TIED_VOCAB=0` restores the flat L7 floor. - **R7: GDN low-band flat route.** Uniform base + recurrent floors (`ssm_out` Q6_K, α/β Q8_0) + full-attention band + stock-parity readout on SSM hybrids. The arch's measured `ffn_down` lever stacks where declared (spark2_5: Q4_K, −0.0183 KLD @ +28 MiB). - **R8: Concave lever stacking.** Joint gains from stacked levers measure at about 91% of the independent additive sum (deep cuts cost super-additively). Validate composite multi-arm builds as a whole. - **R9: Multiplicative head-gate preservation.** Post-projection sigmoid gates compound error along the sequence. The F32 pin costs only ~6 MB. Arch-scoped (spark2_5 correction `gate_pin: F32`). - **R11: Probe-slot economics.** Same-bpw ladder steps (Q4_K/IQ4_NL at 4.5 bpw, Q3_K/IQ3_S at 3.44 bpw) free zero bytes yet consume a probe slot. Deduplicate by byte delta before applying the probe cap; the freed slot goes to a class's depth-1 notch. - **R12: Re-shuffle ride parity, priced (R12-bis).** At or above the anchor's size, the input side re-sells below the anchor tier only at a priced rate. The planner sells only when the probed ride rate beats the best measured buy; an unmeasured ride defaults to parity. - **R13: Dead-buy clamp.** A probed raise that fails the noise floor (ΔKLD ≥ −0.001) records at rate 0. The DP then skips that buy and uses the measured zero rather than the MSE proxy (the mirror of the free-sell clamp). - **R14: Size-band identity (±1.5 pp).** A rung's name is its weight band. An out-of-band build re-homes to the rung that owns that band (recording `planned_tier`) or is withdrawn. The ladder gaps own no rung. `--force-rung` keeps any weight as a measurement product. - **R15: Nearest-fresh-champion targeting** (the rule form of L15). `resolve_anchor` picks the nearest fresh champion above; a recipe receipt routes to the default path. - **R16: Pareto-gain adoption.** A candidate with ΔKLD ≤ −N, concordant value axes, and ΔMiB < 100 is adopted under the aspirational gate. The RMS validator records disagreement up to 2×M and vetoes anything above that. - **R17: Buy-transfer scoping (L18-bis).** Body buys transfer to the verdict regime. Readout-tied buys transfer on dense-tied trunks. Input-side buys do not transfer. SWA band buys transfer only at native ctx. - **R18: Anchor-scoped flattening (L19 refined).** A fine-tune kills buys at the mid band and revives them at the deep band. The detector keys on ≥ 80% dead buys at mid or deep anchors; flat sells separate it from dead-lever families. - **R19: Stock-twin detection and self re-shuffle.** A recipe whose tensor types equal a stock classic carries no Paretrix lever on that topology (byte-identical: Nanbeige Precision ≡ Q6_K-imx, TwIL Precision ≡ Q6_K-imx, ReaderLM Precision ≡ Q6_K-imx, ReaderLM Quality ≡ Q5_K_M-imx). 03 queues one self re-shuffle anchored on the twin at its own size (L17) and duels it at the verdict regime. Measured outcomes: the re-shuffle **wins below the top band** (TwIL Quality 0.0131 → 0.0104, −20.7%, 2.7× the gate; MiniCPM5 Quality −12.5%, Precision −21%) and **ties at the top band** (TwIL Precision +0.0007, inside the gate, so the incumbent holds). The re-shuffle wins or ties by construction. On small qwen35, Precision ≈ Q6_K-imx within the tie zone (Xiaomi +0.0003, not byte-identical). 03 skips the rung by default; `PARETRIX_PRECISION_TWIN=0` forces it. - **R20: Operating-point pricing.** A measured table is a property of its operating point (L14). The default path prices from the anchor nearest the target (≤6 pp is L15 reach; beyond that, the shared lines price), with each class priced in the band of its own anchor floor. Price a Nano from a table measured at its own operating point: class rates escalate ~10× from mid band to dissolution, so the last campaign's top-band table misprices it. ### 5.3 Empirical anchors (selected) | Law / Rule | Anchor | | :--- | :--- | | L1 | qwen35-4B: token_embd Q6_K = 94 MiB per top-p point (93% of Q8_0). | | L2 | ReaderLM-v2 (75% FFN): Compact ffn_down IQ4_XS inverted vs Mini (0.1211 vs 0.1145); Q4_K floor restores 0.0775. | | L4 | Ornith-9B: sub-Q6_K ssm_out → KLD 0.081; flat Q6_K → 0.0241 at Quality footprint. | | L5 | MiniCPM5-1B @ 24%: KLD 0.3886, top-p 66.2%. Sub-3B models close at 24%. | | L6 | lfm2 RCO-36pc 0.0301 vs flat 0.0346 (−13%); nanbeige RCO-33pc 0.0951 vs flat 0.1447 (−34%). | | L7 | TwIL-LM3: token_embd → IQ4_XS = 0.1232; Q6_K restores 0.0500 at the same 27% budget. | | L8 | LFM2.5: protecting 8 sparse attention layers recovers 0.371 → 0.213; shielding mixer params recovers 0.030. | | L11 | MiniCPM5 DSpark 0.4419 (ledger 0.4464), ×1.89 t/s on the 2.6.0 toolchain. | | L13 | qwen35-4B notch rates price a second anchor within 5% (ffn_gate_up 2.60e-4 vs 2.53e-4); RCO-30pc dominates its anchor at equal bytes (0.0385 vs 0.0400). | | L18 | Spark Quality self: +0.0005 @512 tie → −0.0107 @4096 (−18.7% rel). | | L19 | NeoHorse 7/7 dead buys at the Mini anchor; D-gap 34.7% vs base priors. | | R6-ter | Campaign champions at Q5_K/Mini-Nano on ≥15%-share tied trunks (Xiaomi, NeoHorse); Spark (8%) keeps Q6_K. | | R9 | Spark-X2.5: F32 gate pin preserves 0.0044 @30% and 0.0039 @33% for < 6 MB. | | R13 | lfm2 Compact: 3 dead buys priced 0; self re-run recovers −0.0019 (0.0767 → 0.0748). | | R19 | TwIL self re-shuffle: Quality 0.0131 → 0.0104 (won, 2.7× gate) · Precision +0.0007 (tie, incumbent holds). | | R20 | NeoHorse Nano priced from `anchor 'pico'` (27.9%, 0.9 pp); TwIL Nano rejects the 8.6 pp table. | --- ## 6. Precision Guarantees (class floors) | Tensor Class | Floor | Rationale | | :--- | :--- | :--- | | Norms & scales | F16 / F32 | Activation scaling stability. Also a toolchain constraint: `llama-quantize` rejects 1-D tensors and `*_norm.weight`. | | Attention gates | Q6_K / **F32** (spark2_5, R9) | Non-linear routing stability; F32 limits sequential drift. | | Recurrent state (GDN α/β) | Q8_0 | Hidden-state accumulation over long horizons. | | Recurrent readout (`ssm_out`) | IQ4_XS → **Q6_K** (@36%+, L4) | Dynamic range and sequence entropy. | | Readout (untied `lm_head`) | Q6_K (Mini+) / IQ4_XS (low) | Output logit resolution. At most one sell notch (R6-bis top-p fence). | | Token embedding | Q6_K (tied, L7) · R6-ter ladder at share ≥ 15% | Vocabulary representations. | | Speculative heads (MTP/DSpark/DFlash) | Q5_K–Q8_0 | Proposal acceptance rates (L11: drift ≤ 0.03). | | Uncalibrated tensors | IQ4_XS | Fallback for tensors missing from the calibration trace. | | Vision projectors (MMProj) | F32 critical / role-ranked | Spatial visual feature alignment. | **Arch-authored floors** (`corrections.floors.` in the arch JSON) replace the generic ladder and the class hard floor for the classes they name. These are measured knowledge, authored in the JSON with no code change. Current example: qwen35 `gate` (Precision Q6_K · Quality/Compact Q5_K · Mini/Nano Q4_K · Pico/Femto IQ4_XS), the lowest champion tier per rung, distilled from three checkpoints. **Sub-Mini scope.** Pico and Femto default on only for state-rich trunks of 3B or larger. Elsewhere they require an explicit opt-in (`PARETRIX_INCLUDE_PICO=1` / `PARETRIX_INCLUDE_FEMTO=1`). Constructibility gates fire where floors overshoot (Xiaomi: tied vocab at 34% of mass, so Mini-30 is the floor; ReaderLM: tied vocab at the R6-ter boundary). --- ## 7. Evaluation Standard ### 7.1 Protocol `wiki.test.raw`, Flash-Attention, fixed 32,768-token span, BF16 reference logits (per-regime KL bases). **Default = ctx 4096×8** (Long Horizon, the L18 verdict regime). `--medium` 2048×16 · `--light` 1024×32 · `--short` 512×64 (the probe regime, which keeps rates comparable with the generic prior and the noise thresholds). Probes price at 512×64; adoption verdicts bind at 4096×8 in MASD multiples. ### 7.2 Quality bands | Band | KLD | top-p | Tiers | | :--- | :--- | :--- | :--- | | Near-lossless | < 0.0200 | ≥ 95.0% | Fidelity, Precision | | Production | ≤ 0.0850 | ≥ 88.0% | Quality, Compact | | Usable budget | ≤ 0.1650 | ≥ 80.0% | Mini | | Edge service | < 0.4000 | — | Nano, Pico, Femto (≥ 0.40 = entropy dissolution) | The **Pareto column** (11's summary) reads ties at the duel gate `max(MASD, 0.001)`: ★ = frontier · `< X` = strictly dominated by X · `≈ X` = lighter twin inside the tie zone · ≡ = byte-identical. ### 7.3 Verification controls - **Split-integrity ratio check**: recurrent models on partial offload verify PPL against a CPU reference (≤ 1.15× on identical chunks). Suspect results are quarantined. `-ngl` is set explicitly everywhere (0 = CPU). - **Distribution entropy verification**: KLD and top-p are read directly. PPL below the BF16 base flags entropy collapse († in tables), which is a distribution signal rather than a quality gain. - **Calibration-corpus comparability**: a corpus swap shifts absolute KLD by ~0.005 while the rank order holds. Rank products within one corpus, and compare frontier edges across corpora with that shift in mind. - **Calibration-source equivalence**: a Q8_0-sourced imatrix ranks identically to BF16 (ρ 0.999987, top-100 100/100), which makes Q8_0 the default source. - **Artifact-vs-plan verification**: `Paretrix.py verify` compares shipped tensor types against the allocation sidecar. --- ## 8. THE MATRIX — Knowledge Base ### 8.1 File layout ``` arch-paretrix/ ├─ generic-paretrix.json universal prior (bands mid/deep · tails · l19) ├─ -paretrix.json per-arch line: schema paretrix-arch/3 │ facts · corrections · witnesses · model_types · lines{base, finetune} │ lines..sell{class: {segments, esc}} · .buy{class: {depth: {lo,hi,rate}}} │ lines..bands{mid, deep} · .buy_census · .size_mib · .depth_prior └─ … /Paretrix/ ├─ model-paretrix.json schema paretrix-model/2 │ model · arch · model_type · mods · anchors · probes · rates │ allocations · cocktails · champions · notes ├─ eval-cache.json KL results (11) keyed on file × regime × base × binary ├─ paretrix-overhead.json calibrated quant-size overhead factor └─ kld-bf16-ctx.dat(.json) per-regime KL reference bases ``` ### 8.2 Pricing precedence **Model cache (level 1):** the anchor nearest the target (R20), with each class in the band of its anchor floor. **Arch line (level 2):** banded and size-priced. **Generic prior (level 3):** banded and size-priced. Each level overlays the next, class by class. `PARETRIX_FINETUNE=1` forces the L19 modifier (sells ×0.3, buys ×0.5 from `generic-paretrix.json` "l19"). Otherwise, lineage comes from the model cache's `mods`. ### 8.3 The fold (`Paretrix.py matrix`) The fold compiles every model cache under the scan root into the arch lines, per (arch, mods). Each anchor lands at the line's reference size first (`size_mib` = geometric mean of the line's checkpoints). Sells fold banded: each notch falls in the band its anchor floor selects. The fold compiles them conservatively: sells at the range max (non-decreasing with depth), buys at the min of the range lows, and all-dead buy classes at 0.0 (R13, table-wide). The buy census counts measured raises per band (dead = ΔKLD ≥ −0.001) to drive the prior-dead skip. Authored fields (facts, corrections, witnesses, policy) are preserved across every fold, and missing structural facts are detected from a sibling model GGUF. ### 8.4 Analysis tools - **`Paretrix.py atlas`**: rate-field evidence (how tightly each pricing key predicts a measurement), lineage (L19 signature per anchor), probe economy (hindsight skip candidates), routes (what holds each rung), and lint (incoherent tails, non-monotone notches, wide ranges, orphans). - **`Paretrix.py digest [--classics]`**: champion composition per rung (class × MiB × tier counts × E/M/L dominant tier per layer third). This is the rule-extraction view to attach to model cards. - **`Paretrix.py diff [--subset] [--no-kv]`**: byte-level GGUF regression check covering tensor names, types, shapes, payload hashes, and KV pairs (values hashed). --- ## 9. Campaign System (RCO) ### 9.1 Flow ``` anchor (any champion artifact → tensor types → floors) → feasibility preflight (byte arithmetic; closes infeasible rungs in under a second) → notch probes (per class, per depth; prior-dead raises skipped — R13 census) → measured rates (sell segments + buy table; free-sell and dead-buy clamps) → exact-budget DP (anchor floors + measured squeezes + measured raises) → build (llama-quantize with emitted tensor-type rules) → artifact-vs-plan verification → duel at the verdict regime (ctx 4096×8, 1×MASD gate) → adopt (rung name + receipt) · cocktail attempt (auto re-shuffle from arms) → arch fold (shared lines refreshed) ``` ### 9.2 Notch probes The campaign runs one probe per class per depth (`mix-notch-[-dN]`) and one raise per class (`mix-up-`), measured at the probe regime (512×64). Same-bpw steps are deduplicated (R11). Prior-dead raises are skipped without measurement (arch census of ≥2 all-dead, R13). The baseline row replays the anchor allocation as a sanity check. Probe caches are shared between sibling rungs of the same anchor (anchor-stamp keyed). ### 9.3 Exact-budget DP The DP is a multiple-choice knapsack (one option per tied group), solved to the MiB (unit quantization ~15–45 KiB). Squeeze cost per notch comes from the measured segment table (`_kld_squeeze_cost`), and raise gain from the measured buy table. F16 buys fall back to the MSE proxy (L16 hardening). The stage ladder extends 0 (anchor floors only) → 2 → 3 → 4 → 5, and the minimal extension that fits wins. `InfeasibleAlloc` at every stage marks the constructibility gate (family verdict, clean skip). ### 9.4 Duel and adopt Both arms are measured at the verdict regime (4096×8 default, `verdict_ctx` arch override). The gate is 1×MASD of the candidate run (floor 0.001). A gain at or above the gate adopts; anything smaller is a tie, and the incumbent holds. `adopt` writes the verdict into the sidecar (`verdict_kld`, `verdict_vs_kld`) and the receipt into the model cache (one per tier). Verdict guards: a better-recorded incumbent stays in place unless `--force` is given, and a candidate that is not strictly better is refused. ### 9.5 Cocktails Cocktails compose from measured arms. The exchange test runs first (one direction per tensor cluster), then sell bundles fund buys (exact subset search, minimum total cost, ≤8 arms; the ε law). Only value-positive bundles join, and overlap exclusion removes conflicting arms. The additive prediction is probe-regime, because ε conflates composition and regime. Ship a cocktail only on its duel result, since the duel prices ε. --- ## 10. Speculative Modules ### 10.1 MTP 01 extracts the Multi-Token Prediction fusion head (`mtp-BF16.gguf`) through a structural diff between the full and trunk GGUFs. The standalone head has no `token_embd` and cannot load alone; the fused deployment (`14 --mtp`) is the runtime shape. MTP invariance, measured in both directions: PPL and KLD batteries are identical with or without the head attached. ### 10.2 DSpark / DFlash Block-diffusion speculative drafts convert whole (tokenizer borrowed from the parent checkpoint). `Paretrix-modules.py` quantizes them imatrix-free: bridge (fc/eh_proj) ≥ Q6_K · proposal (markov_w/conf_proj) ≥ Q5_K · vocab tables K-quant · unclassified tensors at Q5_K (the F16 catch-all is never used). The trained block size lives in the module metadata (`*.block_size`); 13 reads it, and mask lengths beyond it were never seen in training. ### 10.3 MMProj (vision) CLIP vision towers quantize through a profile grid (`balanced` / `compact` / `fidelity`). Critical tensors (patch and position embeddings, norms, merger) stay pinned at F32, with a deep-boost on trailing blocks. The `llama-quantize` fallback verifies artifact-vs-plan roles; it caught a historical attn/bridge miss with a divergence of ×1.27. ### 10.4 Fusion `14_gguf-module-fusion.py` rebuilds the deployment GGUF (model + MTP head [+ mmproj]). The first file wins on KV. `block_count` widens only on the **fusion base's** key (`.block_count`), and per-layer arrays follow the same prefix. A sub-module's count (e.g. `clip.vision.block_count`) stays untouched. `.part` output renames on completion. --- ## 11. Console Grammar Every script, every depth: | Form | Meaning | | :--- | :--- | | `══ title ═══` | Banner, top-level process | | `── title ───` | Banner, nested process (inherited `PARETRIX_DEPTH`) | | `▸ title` | Section | | ` key value` | Aligned facts | | ` ✓ ok · ✗ failure · ⚠ warning · ⊘ skipped by rule · · note · → command` | Status line | | ` + file size` | Product | | ` ⏳ label · 42.0% (n/N) · 3.1 min · ETA 4.2 min` | One live line per child tool | `PARETRIX_VERBOSE=1` restores full tool logs, tensor listings, and static reading aids. The ETA starts at the first counter sample: model loading is excluded from the speed estimate, and the ETA holds steady between advances. --- ## 12. Environment Variables | Variable | Values | Default | Effect | | :--- | :--- | :---: | :--- | | `PARETRIX_INCLUDE_PICO` / `_FEMTO` | `0` / `1` | scope | 24% / 21% rungs. Default on state-rich ≥3B trunks; `=1` forces, `=0` suppresses. | | `PARETRIX_INCLUDE_PRECISION` / `_FIDELITY` | `0` / `1` | `1` | 42% / 48% rungs. | | `PARETRIX_CLASSIC` | `0` / `1` | `1` | Classic stock line on/off. | | `PARETRIX_CLASSIC_LOW` | `auto` / `on` / `off` | `auto` | Low-band stock quants (Q6_K, Q5_K_M, IQ4_XS, IQ3_M). `auto` enables them on lever-less topologies only. | | `PARETRIX_CLASSIC_EXTRA` | `[,…]` | — | Extra classic formats. | | `PARETRIX_QUALITY_ENGINE` | `native` / `allocate` | `native` | Quality-36 flat vs DP. | | `PARETRIX_PRECISION_GATE_Q8` | `0` / `1` | `0` | GDN sigmoid gates → Q8_0 on Precision. | | `PARETRIX_LOW_FLAT_MIXER` | `0` / `1` | `0` | Low-band flat levers on mixer-dominant hybrids. | | `PARETRIX_FINETUNE` | `0` / `1` | mods | L19 field modifier (sells ×0.3, buys ×0.5 on priors). Auto-detected from the model cache's `mods`. | | `PARETRIX_FEMTO_RECIPE` | `0` / `1` | `0` | Femto class-sell recipe instead of the DP (campaign replication arm). | | `PARETRIX_FEMTO_CLASSIC` | `0` / `1` | `0` | Femto regex-grid measurement arm. | | `PARETRIX_TIED_VOCAB` | `0` / `1` | `1` | R6-ter tied vocab share floors. | | `PARETRIX_ARCH_FLOORS` | `0` / `1` | `1` | Arch-authored floors (`corrections.floors`). | | `PARETRIX_DEPTH_PRIOR` | `` | arch line | Squeeze depth prior override (G5 test arm). | | `PARETRIX_SIZE_NORM` | `0` / `1` | `1` | Size-normalized shared lines (R20). | | `PARETRIX_PRIOR_SKIP` | `0` / `1` | `1` | Prior-dead raise skip (R13 census). | | `PARETRIX_PREFLIGHT` | `0` / `1` | `1` | Campaign feasibility preflight. | | `PARETRIX_PRECISION_TWIN` | `0` / `1` | `1` | When `1`, skips Precision on small qwen35 (R19 C2). | | `PARETRIX_NO_EVAL_CACHE` | `0` / `1` | `0` | Forces re-measurement in 11. | | `PARETRIX_VERBOSE` | `0` / `1` | `0` | Full tool logs and tensor listings. | | `IMATRIX_SOURCE_QUANT` | `` / `bf16` | `Q8_0` | Source format for imatrix generation. | | `PARETRIX_PROBE_NGL` | `auto` / `0` / `` | `auto` | 12's offload for probe evaluations. | | `LLAMA_CPP_DIR` / `LLAMA__PATH` | `` | — | Custom llama.cpp binaries. | --- ## 13. Retest Matrix When changing an engine rule, rebuild only the artifacts the changed condition governs. Compare binary hashes against the archived release: byte-identical artifacts need no re-benchmark, and only differing bytes trigger one. | Rule / Change | Scope | Rungs | Validation target | | :--- | :--- | :--- | :--- | | R6-ter tied vocab floors | Tied trunks, share ≥ 15% | Mini–Femto | Xiaomi-OCR-0, NeoHorse-1-4B | | R6-bis readout parity | Untied, pair > 30% of mass | Nano–Quality | MiniCPM5-1B/2B, PaddleOCR | | R7 GDN flat route | SSM hybrids | Mini, Compact | Ornith-1.5-9B, NeoHorse-1-4B | | R9 F32 gate pin | SWA + sigmoid gates | Mini–Pico | Spark-X2.5-4B | | R12-bis ride parity | Re-shuffles with embd ride | all | MiniCPM5-2B, nanbeige4.2-3B | | R14 size-band identity | All shipped products | all | NeoHorse-1-4B (32.2% out-of-band build → DP rebuild) | | R19 stock-twin self re-shuffle | Dense topologies with inert levers | Quality–Precision | TwIL-LM3-Pro (won/tied), MiniCPM5-2B | | R20 operating-point pricing | Default-path builds | all | NeoHorse (anchor 'mini'), TwIL (table skipped) | | Constructibility gate | Floor-limited families | sub-Nano | Xiaomi-OCR-0 (Mini floor), ReaderLM-v2 (Nano floor) | | L11/L20 draft battery | DSpark/DFlash modules | — | MiniCPM5-2B (0.4419), nanbeige4.2-3B | --- ## 14. Runtime Best Practices - **Streaming GGUF writes** spool tensors through 256 MiB buffers, so peak RAM equals one tensor rather than the whole model. - **One header parse per file per process** (stamp-keyed, released immediately, so Windows rename and delete operations stay safe). Vocab-sized arrays record length only. - **CPU threads** default to physical cores, because SMT siblings contend on shared vector units. - **NTFS Xpress8K** compresses BF16 GGUFs and KL bases (~18%). Quantized payloads and the imatrix stay raw, since their high entropy makes decompression a tax with no gain. - **Eval cache** makes repeated duels and sweeps nearly free (10 min → seconds per model). - **Probe cache sharing** between sibling rungs of the same anchor (anchor-stamp keyed) turns a multi-rung campaign's probe cost into one measurement wave.